Skip to content
Good for Bots

The methodology

What makes a site good for bots?

A readable website gives agents a way in, useful content and clear directions. These are the checks behind your Good for Bots report, with their sources, scoring rules and limits.

Current methodology

How scoring works

The score starts with achievements in two halves: the readable base, scaled to 50 points, asks whether a model can read and follow the HTML our crawler receives; the agent layer, scaled to 50 points, covers what the site offers language models on purpose. Active penalties are then subtracted. The total is rounded to a whole number and cannot fall below zero.

An achievement passes for full credit, warns for half credit and fails for no credit. A non-applicable check is excluded, and the remaining applicable weights in its half determine each check’s share. The maximum contributions below assume every current check applies. Your report records the actual contributions for that scan.

For penalties, pass means no deduction. A warning usually applies half the maximum; some checks calculate an exact deduction from their findings. When one response is a bot challenge, other penalties reading that same refused page defer to the more specific diagnosis. A homepage that gives a crawler without a browser nothing to read (a challenge, an answer other than 200, or an empty page with no Markdown alternative) also holds the score below Fair, at 44 at most, whatever the rest earns.

Experimental and deprecated checks are visible but score nothing. Capabilities never contribute points. Blocking training crawlers is neutral. Refusing our bot in robots.txt ends the scan without a score; an unrecoverably blocked scan also has no score.

This is a bounded technical assessment, not a promise of indexing, inclusion in AI answers or citations. We examine the collected responses, using the homepage and up to three sampled pages for page-level checks, rather than crawling every URL. Free and Pro use exactly the same scoring rules.

  • Poorfrom 0 points
  • Fairfrom 45 points
  • Goodfrom 75 points
  • Excellentfrom 90 points

This page describes the current methodology. Each report retains the rules used for its scan. Specification status is recorded at the review date shown below; a convention or draft is labelled as such.

What earns points

Files, representations and markup that help machines discover and understand your content.

llms.txt

Up to +11.25 points
Active checkAgent layerReviewed

A short Markdown index introduces a website and points a language model towards useful pages.

Why it matters

A curated starting point helps a reader choose relevant material without first exploring menus and navigation.

What we check

We request /llms.txt and read it as Markdown: its H1 title, its length and its link entries, each a list item that starts with a link. Numbered items, bold links and notes after a dash count; lines inside fenced code blocks are examples, not structure. The structural check is separate from the quality check below.

Reading the result

Pass: a usable file has an H1, at least one link entry and at least 100 characters. Warn: the file is shorter than 100 characters, or lacks the H1 or the links. Fail: the file is absent, refused to our crawler, answered by a bot challenge, unreachable or disallowed by robots.txt, or the address answers with an HTML page, with JSON, or with text that has neither an H1 nor a link, as a catch-all route does.

Where this check stops

We only look at the root path, although the convention also permits subpaths such as /docs/llms.txt. Linked pages are not fetched by this check. An H1 alone can satisfy the format but does not meet our usefulness threshold.

Sources & status

Community convention

robots.txt

Up to +3.33 points
Active checkReadable baseReviewed

The robots.txt file communicates which paths crawlers may request. It is an access policy, not a guarantee of indexing.

Why it matters

A readable policy lets a crawler follow the site owner’s instructions before collecting content.

What we check

We parse /robots.txt and check delivery and directives, as RFC 9309 defines them: lines may end in CR, LF or CRLF, and a Crawl-delay or Sitemap line does not end a user-agent group. Whether we may scan the site is decided separately, before scoring. A parseable refusal of our bot's access to the homepage ends the scan without a score, even when served under an unusual HTTP status. A rule that refuses us only a particular file, such as /llms.txt or a sitemap, leaves the scan running: we do not request that file, and the check that needs it treats it as unavailable to crawlers. A robots.txt answered with a bare 401 or 403 is treated as no rules for access, as RFC 9309 allows, but it does not pass this check.

Reading the result

Pass: the file is delivered at HTTP 200 and has usable directives without parse errors. Warn: it declares nothing or has parsing problems, such as a line without a colon, a rule above the first User-agent line, or an empty or unparseable Sitemap URL. Fail: it is missing, empty, unreadable, an HTML page or delivered under a status other than 200.

Where this check stops

Our requirement for HTTP 200, and our partial credit for a file without directives or with lines we cannot parse, are scoring choices; RFC 9309 asks crawlers only to use the rules they can parse. We read a bounded file sample. User-agent names are compared by their letters, hyphens and underscores, so a digit in a name is ignored. This check does not judge whether the declared rules allow other AI crawlers; separate checks do that.

Sitemaps

Up to +8.33 points
Active checkReadable baseReviewed

A sitemap lists public URLs, or points to other sitemaps that do. It provides a machine-readable route into a website.

Why it matters

Discovery does not have to depend on a crawler understanding menus or guessing URL patterns.

What we check

We use the sitemaps declared in robots.txt, or, when none is declared, try /sitemap.xml, /sitemap-index.xml and /sitemap_index.xml. When every declared sitemap answers 404 or with a web page, we also try /sitemap.xml. We read XML and plain-text sitemaps, including gzip files and XML written with a namespace prefix, and sample the sitemaps an index lists. A web page or any other answer that is not a usable sitemap at a location we guessed counts as no sitemap there.

Reading the result

Pass: a sitemap provides usable page URLs, unless every sitemap robots.txt declares is missing. Warn: a sitemap declared in robots.txt cannot be parsed, a sitemap yields no page URLs, for example because the sitemaps its index lists cannot be read, or a declared sitemap cannot be fetched; the report says whether it is missing, disallowed by robots.txt, refused to our crawler, answered by a bot challenge or unreachable. It also warns when every declared sitemap is missing but /sitemap.xml works, because crawlers that follow the declaration find nothing. Fail: robots.txt declares no sitemap and none is found at the conventional locations.

Where this check stops

This is bounded discovery, not a full sitemap audit. We retain at most 200 URLs per document, read at most 2 MB per sitemap and follow indexes one level deep, so a URL count can be a lower bound. When robots.txt declares sitemaps, the only other location we try is /sitemap.xml, and only when every declared one is missing; a declared sitemap that is refused or unreachable is reported as such even if another location works. The default XML namespace, path-prefix restrictions and the availability of every listed page are not validated.

Sources & status

Sitemaps protocol 0.9

llms.txt quality

Up to +7.5 points
Active checkAgent layerReviewed

A useful llms.txt explains what the site and its linked pages contain, rather than providing a bare list of URLs.

Why it matters

A summary and link descriptions give a model context for choosing which source to retrieve.

What we check

We inspect the parsed root llms.txt for a summary blockquote and the proportion of links with descriptions. Any text after a link that says something counts as its description, whether it follows a colon, as the format specifies, or a dash or other separator; the report points out descriptions that do not use the colon. This is a usefulness test on top of the file’s structure.

Reading the result

Pass: a summary and descriptions on at least 60% of links. Warn: at least 25% are described, but the summary or the higher threshold is missing. Fail: no usable file, no links or fewer than 25% described links. The note about separators never changes the result.

Where this check stops

These percentages are our editorial criteria, and accepting separators other than the colon is our reading, broader than the format's. Descriptions are detected structurally; we do not judge their truth, quality or relevance. A missing file fails rather than being excluded from the score.

Sources & status

Good for Bots criterion

Markdown at its own URL

Up to +12.5 points
Active checkAgent layerReviewed

A full-text bundle or linked Markdown pages give agents content at URLs they can discover directly.

Why it matters

An agent can fetch these resources without knowing that a page supports an Accept header preference.

What we check

We read the first 64 KB of /llms-full.txt and count Markdown or MDX links in llms.txt. A bundle and individual page mirrors are alternative ways to satisfy this check. A bundle must be Markdown content: not an HTML page or JSON, with at least one heading or link, and neither a copy of llms.txt nor the homepage the address redirects to.

Reading the result

Pass: a usable bundle of at least 2,000 bytes, or at least three Markdown links in llms.txt. Warn: a smaller bundle or fewer linked mirrors. Fail: neither route is present, and the report says why when the bundle address was refused, answered by a bot challenge, disallowed by robots.txt, unreachable, a copy of llms.txt or a redirect to the homepage; negotiation alone does not pass this check.

Where this check stops

Only a 64 KB sample of the bundle is read, so a copy of llms.txt larger than that is not recognised. Linked mirrors are counted, not fetched, and completeness or equivalence to the website is not verified. Bundle size, link-count and copy thresholds are our criteria.

Sources & status

llms.txt convention and Good for Bots criteria

Markdown negotiation

Up to +10 points
Active checkAgent layerReviewed

Content negotiation lets the same page URL return Markdown to an agent and HTML to a browser.

Why it matters

Markdown can present the text without the navigation, styling and script markup that a reader would otherwise have to remove.

What we check

We request the homepage and up to three pages chosen from the sitemap and homepage links with Accept: text/markdown, and inspect each response's status, media type, body and Vary: Accept header. Each page is graded on its own, and the median page decides. We reject HTML merely labelled as Markdown: a body that starts as an HTML document is HTML whatever its label, while Markdown that mentions HTML tags further in, in a code example for instance, is still Markdown.

Reading the result

Pass: HTTP 200 labelled as Markdown (text/markdown, or the older text/x-markdown), at least 200 characters and Vary: Accept. Warn: useful plain text, a shorter Markdown response or a missing Vary: Accept. Fail: refusal, an error response, HTML or another representation, including the unregistered application/markdown.

Where this check stops

The size threshold is ours, and so is requiring Vary: Accept for a pass: HTTP recommends the header rather than requiring it. We do not establish that the Markdown contains every fact in the HTML. Pages outside the sample are not tested, and a sampled page that could not be read is left out rather than failed. Markdown at a separate URL is assessed separately.

Active checkAgent layerReviewed

A site can address individual AI crawlers in robots.txt instead of relying entirely on a wildcard rule.

Why it matters

Naming crawlers communicates a deliberate policy. This check rewards an explicit position, whether it permits or refuses training.

What we check

We match user-agent tokens against our crawler registry, combine matching groups and record their purposes and access rules. A wildcard group alone does not count as an explicit AI crawler policy, and neither do Content-Signal or Content-Usage lines: they state preferences to every crawler without naming one, and the content signals check scores them.

Reading the result

Pass: at least two crawler purposes are represented, or at least four distinct known crawlers are named. Warn: some are named but neither threshold is met. Fail: no known AI crawlers are named, including when the file states usage preferences for every crawler.

Where this check stops

The thresholds are our criteria, not RFC requirements. Recognition depends on the crawler registry, which we check against the operators' own documentation, and can miss newly introduced names. Earning these points does not mean search crawlers are allowed: that is a separate penalty check.

Sources & status

RFC 9309 and crawler operator documentation

AI usage preferences

Up to +2.5 points
Active checkAgent layerReviewed

Usage preferences describe what may be done with content, separately from which crawlers may fetch it.

Why it matters

An owner can distinguish search, AI use and training instead of treating every use as the same permission.

What we check

We read every Content-Signal and Content-Usage line in robots.txt, including rules scoped to a path, and the Content-Usage header on the final successful homepage response. Findings identify the header’s response URL. Both carriers share one check and count once. Either vocabulary can satisfy the check; recognised positive and negative preferences count equally. Only yes or no, or y or n for Content-Usage, states a preference; any other value states nothing. Content-Usage is read as the IETF draft specifies: parameters after a value are ignored, and a declaration that does not parse, for example one with uppercase keys, states nothing.

Reading the result

Pass: at least one recognised category and value. Warn: a usage declaration exists but its categories or values are not recognised. Fail: no declaration is found.

Where this check stops

HTTP coverage is limited to the final homepage response; its header describes that response, not the whole site. We do not resolve effective permissions across carriers or decide whether a downstream service honours a preference. A rule scoped to a path counts as though it covered the whole site. Drafts and conventions are not all published standards. The check recognises declarations; it is not a legal interpretation of permission.

schema.org structured data

Up to +8.33 points
Active checkReadable baseReviewed

Structured data describes entities such as an organisation, product or article in a form software can read directly.

Why it matters

Explicit names, types and relationships reduce the amount a reader must infer from page layout.

What we check

We inspect JSON-LD in collected server-supplied HTML for readable JSON, a schema.org context, a named subject (a top-level entity, or the main entity or topic of one) and a small set of structural properties. Entities that share an @id are merged first, as JSON-LD requires. Types are compared with every class schema.org defines. We read the homepage and up to three pages chosen from the sitemap and homepage links, grade each page on its own, and let the median page decide.

Reading the result

For each page, pass: the inspected markup meets our usefulness checks. Warn: partial, mixed or incomplete JSON-LD, only types schema.org does not define, or detected microdata/RDFa without JSON-LD. Fail: no structured data or only unreadable blocks. A truncated sample without markup warns, and so does JSON-LD we could not read because a block was too large or was cut off with the page. The median page decides the result.

Where this check stops

The structural property requirements are ours, not schema.org required fields or Google rich-result eligibility. The event and product requirements also apply to their subtypes in schema.org's hierarchy, except event series, course runs, publications, deliveries and the superseded user-interaction types, which carry their dates elsewhere or are not single events. We check type names against the vocabulary, not properties, and do not prove that the claims match visible content. Blocks over 512 KB are not read. JavaScript-injected markup is not read. Pages outside the sample are not inspected.

Semantic HTML

Up to +10 points
Active checkReadable baseReviewed

Semantic markup identifies the main content, headings and page landmarks instead of leaving everything in generic containers.

Why it matters

Content extractors can more easily separate the page’s subject from navigation and other surrounding material.

What we check

We check the raw HTML of the homepage and up to three pages chosen from the sitemap and homepage links, each for one correctly placed visible main region, a heading outline with a named H1 and no downward skips, and at least two other landmark roles. On pages with at least 40 content words, the main region must hold at least half. Content a React server streams is read where React's own script places it. A heading is named by its text, an image's alt text or its label. Words are counted between block elements, and in Chinese, Japanese, Thai and similar scripts by word segmentation rather than spaces. The report also shows the share of extracted text in main, navigation, page header, page footer and other regions, plus the words before main text begins. This diagnostic uses the homepage response and does not change the result.

Reading the result

Each page passes when all three requirements are met, warns when two are, and fails when fewer are met or it supplies no readable HTML content. The median page decides the result.

Where this check stops

The content-share and landmark thresholds are our criteria. This is not an accessibility audit. Stylesheets are not applied. Hidden text counts toward the legacy scoring totals but is excluded from the separate text composition diagnostic where static markup identifies it. Heading levels are read from the tag, without headingoffset. Streaming formats other than React's are not reassembled. Text composition and text-to-HTML ratios do not affect the verdict. Unmarked text is not assumed to be main content, and one page cannot establish repetition across a site. Pages outside the sample are not graded.

Crawlable navigation links

Up to +8.33 points
Active checkReadable baseReviewed

Checks whether navigation in the homepage HTML exposes destinations a reader can follow.

Why it matters

A menu can look complete while exposing no page addresses to a crawler.

What we check

We inspect anchors and declared links inside navigation landmarks and menus. Page links need native hyperlink markup and a usable address. Local section links need a target. Menu-opening controls and contact links are excluded.

Reading the result

Pass means every assessed item is usable. A warning means at least 90% are usable, or the sample is incomplete. Fewer than 90% is a failure. No observed navigation links is not applicable; an unavailable page is a failure.

Where this check stops

This reads collected HTML without executing scripts or loading stylesheets. It does not visit destinations or prove that menus work after interaction. Findings describe the homepage sample.

Sources & status

HTML Living Standard and WAI-ARIA Recommendation; crawlability guidance from Google.

Names for links and buttons

Up to +4.17 points
Active checkReadable baseReviewed

Checks whether links and buttons in the homepage HTML have a name a machine reader can recover.

Why it matters

An unnamed icon button or link gives an agent little evidence of what it does.

What we check

We calculate names from referenced labels, ARIA labels, native text alternatives, element contents and applicable fallbacks. Icon controls count too. A description alone does not supply a name.

Reading the result

Pass means every assessed control is named. A warning means at least 90% are named, or the sample is incomplete. Fewer than 90% is a failure. No observed controls is not applicable; an unavailable page is a failure.

Where this check stops

Names are checked for presence, not meaning. Stylesheets, generated content, scripts and runtime accessibility behaviour are outside this static check. Findings describe the collected homepage.

Sources & status

Accessible Name 1.1 and WAI-ARIA Recommendations, with AccName 1.2 and HTML-AAM draft references.

Active checkReadable baseReviewed

Checks whether observed homepage fields have recoverable names and recognizable input semantics.

Why it matters

An agent needs to identify which field accepts a value and what that field means.

What we check

Native and custom fields are assessed using associated labels and the accessible-name algorithm. Broken explicit form associations are reported. Submit controls and input-type fallbacks are recorded without requiring a particular submission mechanism.

Reading the result

Pass means every assessed item meets the checks. At least 90% correct is a warning; fewer is a failure. Incomplete samples cannot pass. No applicable items on a complete page is not applicable.

Where this check stops

Static markup cannot establish successful submission, validation feedback or correct label meaning. Button names are assessed separately.

Sources & status

HTML Living Standard and W3C accessibility Recommendations; static subset.

Active checkReadable baseReviewed

Checks static control roles, required widget states and references to related elements on the homepage.

Why it matters

State and relationship declarations let an agent distinguish a collapsed control, a selected value and the content a control affects.

What we check

We validate recognized roles, a defined subset of required widget properties, declared state values and ID-reference targets. Native HTML supplies equivalent semantics without redundant ARIA.

Reading the result

Pass means every assessed item meets the checks. At least 90% correct is a warning; fewer is a failure. Incomplete samples cannot pass. No applicable items on a complete page is not applicable.

Where this check stops

This is a static subset of ARIA checks. It does not establish keyboard operation, runtime state accuracy or full accessibility conformance.

Sources & status

HTML Living Standard and W3C accessibility Recommendations; static subset.

Canonical URL

Up to +1.67 points
Active checkReadable baseReviewed

A canonical URL names the one address, among several serving the same content, that should be indexed and cited.

Why it matters

A page is often reachable at several addresses, with tracking parameters or a trailing slash. The canonical tells a search engine or an agent which one to use as the source.

What we check

We read the Link header and the link elements of the homepage and of up to three pages chosen from the sitemap and homepage links. Each page should declare one canonical URL with an http or https address, resolved against the document base the way a browser resolves it. When the target is a page this scan already fetched, we also check that it answered without an error and does not name a different canonical of its own.

Reading the result

A page passes with one usable canonical URL. It warns when the declaration sits in the body instead of the head, when a host name written without a scheme resolves to a nested path, when the target carries a fragment, when a sampled page names the homepage, or when a fetched target answered with an error or names another canonical. It fails with no declaration, with no usable address, or with two different canonical URLs. The median page decides the result.

Where this check stops

Targets outside the sample are not fetched, and a permanent redirect cannot be told from a temporary one. A canonical added only by JavaScript is not read. Findings describe the pages read and cannot establish that every page on the site is canonicalised correctly.

Sources & status

RFC 6596 (Informational); placement and address form follow Google Search guidance.

What can lower the score

Access barriers and missing page basics. Pass means the problem was not found; it does not earn extra points.

Bot challenges

Up to −15 points
Active checkReviewed

A firewall can return an interstitial or challenge instead of the requested content, even when robots.txt permits access.

Why it matters

A client that receives a challenge has no page text to read. Later recovery through a browser transport does not erase the original refusal.

What we check

We read our plain crawler's own requests for the homepage and for its Markdown version, including the page a meta refresh leads to, and look for proof of a challenge: a challenge header from Cloudflare, AWS WAF, Vercel, DataDome or Kasada, markup that only a challenge page carries, a firewall vendor's own block-page wording, or a refusal status of 401, 403 or 418. Generic words such as captcha or access denied count only in the title of an error page, and no wording counts on a long page. Other homepage penalties defer when the same refused response belongs to this diagnosis.

Reading the result

Fail: the direct homepage request meets a recognised challenge. Warn: only the request asking for Markdown is challenged. Pass: no challenge is recognised. N/A: the homepage never supplied a response we can assess.

Where this check stops

Detection uses known signatures and can miss an unfamiliar challenge, particularly one served with HTTP 200. A 401 or 403 on the homepage counts even without a vendor's marker, because a crawler gets no page either way, so a homepage behind a login is charged here too. A plain rate limit (HTTP 429 with no vendor's marker) ends the scan without a score instead of counting here. A refusal in robots.txt is an opt-out, not a challenge penalty; an unrecoverably blocked scan may have no score at all.

Active checkReviewed

A page can ask crawlers not to index it, or to withhold text snippets, through metadata and HTTP headers.

Why it matters

These instructions restrict discovery and reuse even when the page can be fetched successfully.

What we check

We inspect homepage X-Robots-Tag headers and recognised robots meta directives. A directive counts when it applies to all crawlers or is addressed to a search engine such as googlebot or bingbot. One addressed to an AI training crawler, such as GPTBot, is recorded as a training opt-out without a penalty, as are training-only opt-outs such as noai; one addressed to any other crawler is ignored.

Reading the result

Fail: noindex or none is present for all crawlers or for a search engine. Warn: nosnippet or max-snippet:0 for the same crawlers, without a full indexing block. Pass: neither restriction is found. N/A: availability or a bot challenge prevents assessment of the real homepage.

Where this check stops

Only the homepage is assessed. A directive addressed to one search engine counts as much as one for all crawlers, although other crawlers may behave differently, and a search engine outside the names we recognise is ignored. With preserved header lines, a crawler name applies across commas until another crawler name or a new header line. Older evidence or a transport without those boundaries uses a conservative fallback: an unprefixed directive after a comma counts for all crawlers. Boundaries already merged by an intermediary cannot be recovered. We do not claim to know whether a search engine has indexed the page.

Content without JavaScript

Up to −25 points
Active checkReviewed

An empty HTML response with executable scripts can leave readers that do not run JavaScript with nothing to read.

Why it matters

Content in the server response is available without downloading and running an application.

What we check

We look for a complete, textless homepage response containing executable scripts. Short pages, nonempty noscript fallbacks, and textless pages containing media, controls or other potentially meaningful elements remain undecided. We also look for negotiated Markdown or a real llms-full.txt. No scripts run during this check.

Reading the result

Fail: an empty script-bearing response without an observed Markdown alternative. Warn: the same response with an alternative. Pass: at least 120 content words were captured. N/A: incomplete or unavailable HTML, or insufficient evidence to establish an empty script-bearing page. Short pages receive no penalty.

Where this check stops

This conservative check misses partial hydration and shells containing loading messages, fallback notices or graphics. Script presence does not prove what the script does. A pass does not establish complete content, and a Markdown alternative is not checked for equivalence.

Sources & status

Good for Bots heuristic; no external specification

Search crawler access

Up to −25 points
Active checkReviewed

Search and user-fetch crawlers have a different purpose from crawlers collecting training data.

Why it matters

Blocking a discovery or user-requested retrieval route declares that the site does not want to be read for AI answers, and crawlers that honour robots.txt stay away. Some operators may still fetch a page when a person asks for it. Refusing training remains neutral in our score.

What we check

We apply robots.txt rules to the homepage path for the search and user-fetch crawler tokens their operators document. Explicit groups, wildcard fallback and rule precedence determine access; training, agent and extraction tokens are not charged. Only a robots.txt served normally, with HTTP 200, is read for this check.

Reading the result

Pass: no applicable rules block the selected crawlers. Fail: one or more are blocked, with the exact deduction recorded in the report. N/A: the robots probe was not run. A missing robots file does not itself mean access is blocked.

Where this check stops

This evaluates declared access to the root path, not live requests impersonating other bots. Firewall behaviour and deeper paths can differ. OpenAI, Perplexity, Meta, Amazon and Google say their user-triggered fetchers may fetch a page robots.txt refuses them when a person asks for it, so we charge the declared refusal, not proven invisibility. Recognition depends on our maintained crawler registry, and we charge only tokens their operators document.

Homepage availability

Up to −10 points
Active checkReviewed

The homepage needs to be reachable and delivered securely before a crawler can use it as a source.

Why it matters

An error response or an insecure final address undermines the entry point that other discovery signals depend on.

What we check

We inspect the final homepage response and URL, and whether our plain HTTP request, the kind a crawler without a browser makes, received a response. An HTTP entry point that redirects to HTTPS is accepted. A recognised challenge is assessed by the bot-challenge check instead. Every scan starts over HTTPS, so a site that answers only plain HTTP cannot be scanned and gets no score. A homepage that answers HTTP 451 or redirects to another domain also ends the scan without a score.

Reading the result

Pass: the final response is HTTP 200 over HTTPS, and our plain request received a response. Fail: no response, another final status, an HTTP final URL, or a page we reached only through a browser because our plain request got no response, even when retried. N/A: a recognised challenge owns the refusal.

Where this check stops

Exactly HTTP 200 is our delivery criterion, not a claim that other successful HTTP statuses violate the standard. We do not perform a separate TLS configuration audit or check availability across regions. We scan from the EU, and a page served there with HTTP 200 is graded as the homepage, even when it only says the site is unavailable in the region.

Sources & status

HTTP RFC 9110 and Good for Bots criteria

Active checkReviewed

The page title and meta description give a short identity and summary for the homepage.

Why it matters

They help readers and discovery tools understand what the page is about before extracting its full content.

What we check

We look for non-empty title and meta name=description values in the head of the server-supplied HTML. Social-card metadata does not substitute for either field.

Reading the result

Pass: both fields are present and non-empty. Fail: either is absent or empty. N/A: no successful HTML homepage is available, or a challenge owns the response.

Where this check stops

We do not grade wording, length, duplicate values or factual accuracy. Metadata inserted only by JavaScript is not read. An HTML parser can repair misplaced elements before this check sees them.

Sources & status

HTML Living Standard and Good for Bots criteria

Redirect chains

Up to −3 points
Active checkReviewed

This report used different scoring rules from the current methodology. The report’s recorded score has not been recalculated.

A redirect chain is the sequence of additional requests before the homepage reaches its final address.

Why it matters

Each extra hop adds work and delay before an agent can read the content.

What we check

We count followed HTTP redirects and the same-site declarative refresh supported by our fetcher. Only chains that successfully reach an unchallenged HTTP 200 homepage are graded.

Reading the result

Pass: at most two hops. Fail: more than two hops. N/A: the chain does not yield a gradable homepage, so availability or challenge detection handles the failure.

Where this check stops

Two hops is our threshold, not an RFC limit. The stored evidence records the count rather than every intermediate redirect rule. Unsupported JavaScript navigation is not followed, and different transports may encounter different chains.

Sources & status

HTTP RFC 9110 and a Good for Bots threshold

Crawler response time

Up to −3 points
Active checkReviewed

This measures the time our ordinary HTTP crawler spends obtaining the homepage response.

Why it matters

A slow response consumes a reader’s time budget before content processing can begin.

What we check

We use the direct request’s elapsed time through reading the capped body, including connection setup and redirects but excluding our own pacing waits. Browser transport startup is not part of the measurement. When a request times out, loses its connection or meets a gateway error, we ask once more and time the second attempt, not the failed one.

Reading the result

Pass: a direct, unchallenged HTTP 200 response completes within 3,000 ms, including exactly 3,000 ms. Fail: it takes longer. N/A: a direct error, challenge or non-200 response prevents assessment.

Where this check stops

Three seconds is our threshold. This is not TTFB, LCP or Core Web Vitals. One observation from one region is sensitive to network conditions, and oversized pages are timed only through the retained sample. The timer is elapsed time on our side, which can include a brief delay of our own while other scans run; that matters only close to the threshold.

Sources & status

Good for Bots threshold; HTTP semantics from RFC 9110

Capabilities for agents

Discovery signals for services and tools. A passing check adds a capability to the badge; absence never costs points.

API Catalog

Badge only · no score impact
Active checkReviewed

An API catalogue advertises APIs and related machine-readable descriptions in a linkset document.

Why it matters

It gives an agent a discovery route into services the site makes available to external clients.

What we check

We request /.well-known/api-catalog and can follow the first catalogue advertised by a homepage typed link. The returned JSON and media type are inspected; a successful HTTP status alone is insufficient. A web page or other document that is not a linkset is read as no catalogue, the way a site answers a path it does not implement, unless it is served with the linkset media type.

Reading the result

Pass: a usable linkset with at least one link. Warn: a document served as a linkset that is not a valid one, a linkset whose entries carry no links, or an advertised catalogue that could not be read. N/A: no catalogue is established, including a web page or unrelated document at the path, and unreadable challenged or truncated evidence.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. We accept valid linksets without insisting on the item relation or the specified media type. Links inside the catalogue are not followed, so endpoint availability is not established. Only bounded homepage discovery is performed. Absence never costs points.

Sources & status

RFC 9727 and RFC 9264

auth.md

Badge only · no score impact
Experimental checkReviewed

An auth.md document explains how an agent can discover authentication or registration for a service.

Why it matters

An explicit guide can replace guesswork about credentials and the next discovery step.

What we check

We read /auth.md and distinguish agent-facing authentication guidance from an ordinary product page. Our passing signal is a reference to OAuth protected-resource or authorization-server well-known metadata.

Reading the result

Pass: a recognised document points to the OAuth discovery metadata. Warn: agent-facing authentication guidance exists without that reference, including guidance saying registration is not supported. N/A: no relevant document is recognised.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. auth.md is a reference implementation rather than a conformance standard. We do not require its template headings, follow the metadata links, register an account or test authentication. Lexical recognition is English-biased.

MCP Server Card

Badge only · no score impact
Experimental checkReviewed

An MCP server card advertises a server and the remote connection routes it offers.

Why it matters

A discovery document lets an agent find a declared endpoint before starting an MCP session.

What we check

We read MCP card entries in /.well-known/ai-catalog.json, either inline or through a bounded linked-card fetch. A card needs a name, description and a supported streamable-http or sse remote.

Reading the result

Pass: at least one card meets our discovery criteria. Warn: a declared card is incomplete, unusable or could not be retrieved. N/A: no card is established, or a listed card was not fetched within this scan.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. This is a usability threshold, not full schema conformance. We do not contact the MCP endpoint. Schema, naming and hosting deviations are notes. Legacy /.well-known/mcp.json and /.well-known/mcp/server-card.json paths are not probed; nested catalogues are not explored.

Sources & status

Experimental MCP extension and AI Catalog draft

A2A Agent Card

Badge only · no score impact
Experimental checkReviewed

An A2A card describes an agent and the protocol interfaces through which another agent can contact it.

Why it matters

Declared bindings distinguish an A2A service from a generic URL or an MCP server.

What we check

We read /.well-known/agent-card.json and inline or linked A2A entries in the AI Catalog. A usable card needs a name, description and a JSONRPC, GRPC, HTTP+JSON or absolute-URI binding at a suitable address other than the card itself.

Reading the result

Pass: a card advertises a usable A2A interface. Warn: a recognised card is incomplete or advertises only incompatible bindings such as MCP. N/A: no card is established or the required evidence was not collected.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. The legacy /.well-known/agent.json path is not probed. Older card shapes can still be read at supported discovery locations. Endpoints and signatures are not tested; a declared A2A binding can point at a service that does not implement it.

Sources & status

A2A protocol and AI Catalog draft

Agent Skills discovery

Badge only · no score impact
Experimental checkReviewed

A skills index advertises instructions or packaged skills that an agent can discover and retrieve later.

Why it matters

Discovery metadata makes the available skills visible without downloading or executing every artifact.

What we check

We inspect both /.well-known/agent-skills/index.json and /.well-known/skills/index.json, plus typed AI Catalog entries. We recognise v0.2.0 and legacy v0.1 indexes, requiring at least one usable entry. v0.2.0 needs a name, description, distribution type, artifact URL and syntactically valid SHA-256 digest; v0.1 needs a name, description and a file list containing SKILL.md.

Reading the result

Pass: a complete index contains a usable entry. Warn: an advertised index is malformed, uses an unknown schema or has no usable entries. N/A: absent, empty, unrelated or unreadable discovery evidence. Valid entries can survive invalid siblings.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. Artifacts are not downloaded or executed. Digests are checked for syntax, not against file bytes. Only one distinct linked index is fetched when needed, so later entries can be missed. The badge establishes discovery, not safety, integrity or successful use.

WebMCP declarations

Badge only · no score impact
Experimental checkReviewed

WebMCP lets a page declare tools for browser-based agents through forms or JavaScript registration.

Why it matters

Tool names and descriptions can expose actions more directly than asking an agent to infer them from interface layout.

What we check

We inspect collected HTML for complete static form declarations or inline modelContext.registerTool calls with readable names, descriptions and callbacks. The current document API and legacy navigator spelling are recognised.

Reading the result

Pass: a complete static declaration appears on a non-truncated HTTPS page. Warn: markers exist but declarations are incomplete, duplicated, dynamic, insecure or cut off. N/A: no markers or no usable HTML are available.

Where this check stops

No tools are executed and no forms submitted. External scripts are not downloaded. Browser support, policy, callback behaviour and schema validity are not verified; a declaration in an uncalled branch can pass. This is discovery rather than proof of a working tool.

Sources & status

W3C Community Group draft and an API explainer

Web Bot Auth

Badge only · no score impact
Experimental checkReviewed

A Web Bot Auth key directory publishes the public keys a site's own bots and agents use to sign their HTTP requests.

Why it matters

Websites receiving those requests can verify that they come from this operator, instead of trusting a User-Agent string anyone can copy.

What we check

We request /.well-known/http-message-signatures-directory and read it as a JSON Web Key Set. A key must be a supported public signing key whose kid, if present, matches its thumbprint, and it must not be a published test key. When the response is signed with the directory tag, we verify the signature and its content digest.

Reading the result

Pass: at least one usable signing key. Warn: private key material, a document that is not a key set, plain HTTP, a redirect to another host, no usable keys, or a response signature that fails. N/A: no directory, an unrelated or empty document, or unreadable evidence.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. We cannot see the requests a site's bots send elsewhere, so the badge does not show the keys are in use. Keys published outside the well-known directory are not found. An unsigned directory can pass, although Cloudflare's verifier requires a signature.

x402 and MPP payments

Badge only · no score impact
Experimental checkReviewed

x402 and the Machine Payments Protocol let a site answer an unpaid request with HTTP 402 and a machine-readable payment request.

Why it matters

An agent that meets a payment request it can read can pay for one request and continue, instead of being turned away or needing an account.

What we check

We read any HTTP 402 on the homepage and robots.txt, request /.well-known/x402, and fetch one resource it lists with GET. An x402 PAYMENT-REQUIRED header or version 1 body, or an MPP Payment challenge in WWW-Authenticate, must carry what a client needs to construct a payment.

Reading the result

Pass: a usable x402 or MPP payment challenge. Warn: a challenge that cannot be used, a listed resource that answers without one, or a manifest listing nothing we can request. N/A: no 402 and no manifest, a 402 using neither protocol, or a facilitator-only manifest.

Where this check stops

When robots.txt disallows a discovery URL, we report that it could not be checked; this does not prove the document is absent. Paid endpoints the site does not list, or lists only for POST, cannot be found. We never pay, so whether a payment would succeed is not verified. One listed resource is requested per scan.

ACP discovery

Badge only · no score impact
Experimental checkReviewed

ACP lets a business publish how agents can access its commerce integration.

Why it matters

Explicit discovery metadata gives an agent a starting point without guessing API addresses or creating a shopping session.

What we check

We fetch the protocol's discovery document and check its identity, versions and usable service declarations. The same validated evidence contributes to the site's API and commerce profile.

Reading the result

Pass: usable discovery metadata is published. Warn: recognisable metadata is incomplete or unusable. N/A: no declaration is established, or the document could not be read. Absence never costs points.

Where this check stops

This checks published declarations, not a completed purchase or full protocol conformance. Endpoints, referenced schemas, payment configuration and authentication are not exercised. Integrations published only by a provider elsewhere may remain undetected.

Sources & status

Open commerce protocol specification; experimental scanner capability

UCP discovery

Badge only · no score impact
Experimental checkReviewed

UCP lets a business publish how agents can access its commerce integration.

Why it matters

Explicit discovery metadata gives an agent a starting point without guessing API addresses or creating a shopping session.

What we check

We fetch the protocol's discovery document and check its identity, versions and usable service declarations. The same validated evidence contributes to the site's API and commerce profile.

Reading the result

Pass: usable discovery metadata is published. Warn: recognisable metadata is incomplete or unusable. N/A: no declaration is established, or the document could not be read. Absence never costs points.

Where this check stops

This checks published declarations, not a completed purchase or full protocol conformance. Endpoints, referenced schemas, payment configuration and authentication are not exercised. Integrations published only by a provider elsewhere may remain undetected.

Sources & status

Open commerce protocol specification; experimental scanner capability

See what your website gives an agent.

Your report connects these criteria to the evidence collected from your site.

Check your website →About the score badge