Skip to content
Good for Bots

Practical guide · Page content

How to serve content AI crawlers can read without JavaScript

Which AI crawlers run JavaScript, how to fix empty app shells in Next.js, Nuxt, SvelteKit or Vite, and how to test the raw HTML response yourself.

By Good for Bots · Reviewed

Read as Markdown ↗

Put the text you want read into the HTML your server sends: render it on the server, generate it at build time or prerender it. Many crawlers that collect pages for AI products parse that HTML without running JavaScript, so content that appears only after a script runs is not there for them. Keep scripts for interactivity, and make sure the words are already in the response.

A browser hides the difference. It downloads the scripts, runs them and paints the result, so a visitor sees the same page either way. This guide shows how to tell what your server sends, how to fix it in common frameworks and how to check the result. Readable HTML does not guarantee that an assistant will fetch, use or cite your pages; it removes one reason it cannot.

Start with a complete example

The response below is the homepage of Harbour, a fictional scheduling service. The paths are placeholders.

<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Harbour: scheduling for teams that work in shifts</title>
    <meta name="description" content="Harbour builds rotas, handles shift swaps and checks holiday requests against the cover each role needs.">
    <link rel="stylesheet" href="/assets/app-3f9c1e.css">
    <script type="module" src="/assets/app-3f9c1e.js"></script>
  </head>
  <body>
    <header>
      <a href="/">Harbour</a>
      <nav aria-label="Main">
        <a href="/features">Features</a>
        <a href="/pricing">Pricing</a>
        <a href="/docs">Documentation</a>
      </nav>
    </header>
    <main>
      <h1>Scheduling for teams that work in shifts</h1>
      <p>Harbour is a fictional scheduling service. A manager builds the rota once, Harbour repeats it, handles shift swaps and warns when a shift is left without cover.</p>
      <h2>What Harbour does</h2>
      <ul>
        <li>Recurring rotas with weekly, fortnightly or custom patterns.</li>
        <li>Shift swaps that a colleague accepts and a manager approves.</li>
        <li>Holiday requests checked against the minimum cover you set for each role.</li>
      </ul>
      <h2>Who uses it</h2>
      <p>Clinics, cafés, warehouses and support desks with between 5 and 500 people. Each workspace has its own time zone, so a team spread across two countries sees every shift in local time.</p>
      <h2>Plans</h2>
      <p>The free plan covers one team of up to 10 people. Paid plans are billed monthly for each active member; the pricing page lists the current rates.</p>
      <div id="cover-calculator">
        <p>The cover calculator estimates how many people each shift needs. It runs in your browser; everything above works without it.</p>
      </div>
    </main>
    <footer>
      <a href="/support">Support</a>
      <a href="/privacy">Privacy</a>
    </footer>
  </body>
</html>

The heading, the product description, the plans and the links are in the markup. The module script in the head still loads and makes the page interactive in a browser, a step called hydration. The cover calculator is the one part that needs JavaScript, so its container already says what it is rather than waiting empty. This response passes our content without JavaScript check at the version reviewed for this guide.

Which crawlers run JavaScript

Googlebot does. Google describes three phases, crawling, rendering and indexing, and says a page waiting to be rendered "may stay on this queue for a few seconds, but it can take longer than that." Apple says Applebot "may render the content of your website within a browser". Bing announced in 2019 that it uses Microsoft Edge "to run JavaScript and render web pages". Its current guidelines ask sites to avoid "hiding critical content behind client-side rendering" and warn that "content that cannot be reliably rendered may not be indexed or selected for grounding results." Google recommends the server-side route as well: "server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript."

The operators of AI crawlers mostly do not say. The crawler pages of OpenAI, Anthropic, Perplexity and Meta do not mention JavaScript. Common Crawl, whose archive is often used as training data, states that "JavaScript is not executed". Anthropic's documentation for the web fetch tool in its API, which developers give to their own agents, says it "does not support websites dynamically rendered with JavaScript". That tool is not ClaudeBot, and the statement says nothing about the crawler.

The largest published measurement comes from Vercel and MERJ, who monitored crawler traffic on nextjs.org and Vercel's network and published the results on 17 December 2024. They found that "none of the major AI crawlers currently render JavaScript", naming crawlers from OpenAI, Anthropic, Meta, ByteDance and Perplexity. OpenAI's and Anthropic's crawlers did download JavaScript files, in 11.50% and 23.84% of their requests, without executing them. Gemini used Googlebot's infrastructure and rendered pages.

Those figures describe one network in 2024, and an operator can change its crawler at any time. Our scanner does not try to model each crawler. It reads the HTML response as a crawler that does not render would, and never runs the page's scripts, even when it has to fetch through a browser. If the content is in the response, it does not matter which crawlers render.

Why an empty shell reads as an empty page

A client-side application often sends a document like this and builds the page in the browser:

<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Harbour: scheduling for teams that work in shifts</title>
    <meta name="description" content="Harbour builds rotas, handles shift swaps and checks holiday requests against the cover each role needs.">
    <script type="module" crossorigin src="/assets/index-8f2a1c.js"></script>
    <link rel="stylesheet" crossorigin href="/assets/index-4b7d0e.css">
  </head>
  <body>
    <div id="root"></div>
  </body>
</html>

The title and description are present, but the body holds one empty div. A crawler that parses this without running index-8f2a1c.js gets a title and nothing else: no product description, no headings and no links to the site's other pages. Our check fails this response when the site offers no Markdown alternative.

Choose how the HTML gets its content

web.dev's Rendering on the Web describes the options:

  • Server-side rendering: "Rendering an app on the server to send HTML, rather than JavaScript, to the client." It suits pages that change per request or often.
  • Static rendering "happens at build time", producing "a separate HTML file for each URL ahead of time." It suits content that changes when you publish.
  • Prerendering: "Running a client-side application at build time to capture its initial state as static HTML." It can retrofit an existing client-side application.

Any of the three puts the content in the response. Streaming server rendering sends the HTML "in chunks", but within the same response, so a crawler that reads the whole response receives all of it. Hydration then adds behaviour to HTML that already has its content.

Avoid a prerendered copy served only to crawlers you recognise by their user agent. Google calls this dynamic rendering "a workaround and not a long-term solution" and recommends "server-side rendering, static rendering, or hydration" instead. Every reader missing from your list, including a new crawler, an agent or our scanner, still receives the shell.

Fix it in your framework

Most of the frameworks below render on the server or at build time unless told otherwise. An empty shell usually comes from a setting that turns this off, or from a starting point that renders only in the browser.

  • Next.js, App Router. Pages and layouts are Server Components by default, and Client Components are prerendered into the HTML as well, so "use client" on its own does not empty a page. Look for three other causes. dynamic(() => import(...), { ssr: false }) disables prerendering for that component. In a prerendered route, useSearchParams makes the Client Component tree "up to the closest Suspense boundary" render in the browser; wrap the component that calls it in <Suspense>, as Next.js recommends, and everything above it stays in the HTML. Data fetched in useEffect arrives after the HTML; fetch it in a Server Component.
  • React with Vite. The React template that create-vite generates is a single-page application: the body of its index.html is an empty <div id="root"> and a module script. Vite's server-rendering API is, in its own words, "a low-level API meant for library and framework authors". React's documentation recommends "starting with a framework", and the React team deprecated Create React App for new apps in February 2025. For an existing app, move to a framework that renders on the server or at build time, or prerender its routes during the build.
  • React Router, framework mode. "Server rendering is enabled by default." Setting ssr: false switches to SPA Mode, which prerenders only the root route into index.html. To keep static hosting, list your URLs in the prerender option of react-router.config.ts instead.
  • Nuxt. Universal rendering is the default, and Nuxt notes that "as the content is already present in the HTML document, crawlers can index it without overhead." ssr: false in nuxt.config.ts makes the whole app client-side; prerendered that way, it produces "HTML pages with an empty <div id="__nuxt"></div>". Remove the setting and wrap only the parts that cannot render on the server in <ClientOnly>, as Nuxt suggests.
  • SvelteKit. "By default, SvelteKit will render (or prerender) any component first on the server". With ssr set to false, "it renders an empty 'shell' page instead." Look for export const ssr = false in a +layout.js as well as in pages: in the root layout it turns the entire app into a single-page application.
  • Astro. Astro renders components to HTML by default, and components with client:load or client:visible are rendered to HTML before they hydrate. client:only "skips HTML server rendering". A server island marked server:defer is requested from the browser after the page loads, so a crawler reads only its fallback content.
  • Angular. Angular renders in the browser by default. Create a project with ng new --ssr, or add server rendering to an existing one with ng add @angular/ssr. Angular then "prerenders your entire application and generates a server file" by default.

On a hosted site builder, look for its server-rendering or prerendering setting, then check the published response as described in the next section.

Check what your server sends

Look at the response, not at the page your browser paints.

  1. Turn off JavaScript. In Chrome DevTools, open the Command Menu with Control+Shift+P, or Command+Shift+P on a Mac, run Disable JavaScript and reload. Chrome keeps JavaScript off in that tab "so long as you have DevTools open." What remains is close to what a non-rendering crawler parses, although your stylesheet still decides what you can see.
  2. Read the source. Open view-source: followed by the address, and search for a sentence from the middle of the page. A match inside an ordinary element counts. A match inside a <script> block does not: frameworks often embed a page's data there as JSON, and a parser that skips scripts never reads it.
  3. Fetch it without a browser. Save the response with our crawler's user agent and search it the same way:
curl --silent --location \
  --user-agent 'GoodForBotsBot/1.0 (+https://goodforbots.com/bot)' \
  https://your-domain.example/ > homepage.html

Do not delete every line containing <script> to tidy the file: a minified page can be a single line, and you would delete the whole document. Repeat the check on a product page, a documentation page and an article, not only the homepage. Our check reads the homepage alone.

Fix common mistakes

Content fetched after the page loads

<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Harbour: scheduling for teams that work in shifts</title>
    <script type="module" src="/assets/app-3f9c1e.js"></script>
  </head>
  <body>
    <header>
      <a href="/">Harbour</a>
      <nav aria-label="Main">
        <a href="/features">Features</a>
        <a href="/pricing">Pricing</a>
        <a href="/docs">Documentation</a>
      </nav>
    </header>
    <main>
      <p>Loading…</p>
    </main>
    <footer>
      <a href="/support">Support</a>
      <a href="/privacy">Privacy</a>
    </footer>
  </body>
</html>

The layout is server-rendered and the content is not: the script asks an API for it after the page arrives. A crawler reads the navigation, the footer and "Loading…". Our check leaves this response undecided rather than failing it, because a short page with a little text can be genuine. Fetch the data on the server and render it into the HTML; in Next.js, fetch it in a Server Component instead of in useEffect.

A noscript message instead of content

The HTML standard says <noscript> "represents nothing if scripting is enabled, and represents its children if scripting is disabled." A message such as "Please enable JavaScript" is therefore all a non-rendering reader gets from an otherwise empty page. The standard calls the element "a blunt instrument" and suggests the reverse: a page that works without scripts, which the script then turns into the interactive version. Keep <noscript> for a part that cannot work without scripts, such as a link to a plain version of a calculator. Our check leaves pages with non-empty <noscript> content undecided.

Text that appears only after a click

Some tab, accordion and "read more" components create their text only when opened. Render every panel into the HTML and let the script show and hide them. For a single expandable section, <details> and <summary> work without scripts: the standard defines details as "a disclosure widget from which the user can obtain additional information or controls." It also says tab widgets "are not disclosure widgets", so do not build tabs from it.

Every address answers with the same shell

A static host set up for a client-side application answers every path with index.html. Each URL then returns the same empty document with HTTP 200, including addresses that do not exist. Generate an HTML file for each route at build time, or render on the server, and return 404 for missing pages. The same fallback can answer for /robots.txt and /llms.txt; serve those files ahead of it.

Add Markdown as well, not instead

A Markdown version of a page, served when a client asks for Accept: text/markdown, or a complete /llms-full.txt, gives agents that look for it the content without a browser. It does nothing for a crawler that requests the HTML, so our check treats an empty shell with such an alternative as only partly fixed; the current criteria are below. An llms.txt on its own is an index, and its links lead back to the pages that need JavaScript. How to write a useful llms.txt covers the index.

Keep it working after changes

A dependency upgrade, a layout wrapper that renders nothing until it has mounted in the browser, or one ssr: false can turn server-rendered pages back into shells with no visible change in the browser. Add a step to your deployment: fetch a few important pages and fail if a sentence you expect is missing from the response, outside any <script>.

Scan your site to see the HTML our crawler received from your homepage and how the check read it.

Sources and review

Reviewed on 25 September 2026 against these sources:

The examples are ours, and each is tested against the check listed below at the version reviewed for this guide.

Current scanner criteria

These criteria come from our current standards catalogue. They describe what Good for Bots checks, including usefulness rules of our own that the format does not require, and what our detection cannot see. Each report keeps the methodology of the scan that produced it.

Content without JavaScript

active check · Reviewed

We look for a complete, textless homepage response containing executable scripts. Short pages, nonempty noscript fallbacks, and textless pages containing media, controls or other potentially meaningful elements remain undecided. We also look for negotiated Markdown or a real llms-full.txt. No scripts run during this check.

Results: Fail: an empty script-bearing response without an observed Markdown alternative. Warn: the same response with an alternative. Pass: at least 120 content words were captured. N/A: incomplete or unavailable HTML, or insufficient evidence to establish an empty script-bearing page. Short pages receive no penalty.

Limitations: This conservative check misses partial hydration and shells containing loading messages, fallback notices or graphics. Script presence does not prove what the script does. A pass does not establish complete content, and a Markdown alternative is not checked for equivalence.

Full methodology and sources →

Check your site →

Crawler access

How to write robots.txt rules for AI crawlers

Write robots.txt rules that treat AI training, search and user-triggered crawlers separately, with tested examples, the current token list and common fixes.

Reviewed

Discovery files

How to write a useful llms.txt

Write an llms.txt that tells AI agents what your site covers: a complete tested example, the format explained, publishing checks and fixes for common mistakes.

Reviewed

All guides →