Prerendering a React SPA for crawlers
How a page that renders perfectly in a browser can serve a crawler a blank shell, and the build-time fix that makes a single-page app fully indexable.
By Andrew J. Pyle
I write my essays into a database and read them back on a React front end. In a browser they look perfect. So it took me embarrassingly long to check the one view that matters for search: the one a crawler sees when it does not run your JavaScript. Every single essay was serving a blank.
This is the bug that leaves no trace in the browser, why it hit my essays but not my project pages, and the build-time prerender fix that made them crawlable.
01THE SYMPTOM
Two pages, one URL
I write my long-form essays into a database and read them back on a React front end. In a browser they look perfect: title, body, table of contents. So it took me embarrassingly long to check the one view that matters for search: the one a crawler sees when it does not run your JavaScript.
Every single essay was serving a blank. The check is one command. Ask for the page the way a crawler does, with scripts off, and read the title it gets back.
# Fetch the raw HTML a non-JS crawler receives,
# and read just the <title> it hands back.
curl -s https://your-site.com/some-page | grep -o '<title>[^<]*</title>'02WHY IT HAPPENS
The crawler skips the code that loads your content
The page fetches its own body. The component mounts, and inside an effect it calls the API for the content. That effect only runs in a real browser that executes JavaScript. A crawler that reads the raw HTML never runs it, so it receives the empty shell that was there before the fetch.
03THE TELL
Static bodies render, fetched bodies do not
The confusing part was that some pages on the same site, built by the same prerender, were perfectly crawlable. The difference is where the content lives.
A page whose body is static data inside the component renders into the HTML at build time. A page whose body is fetched at runtime does not. Same site, same pipeline, two different outcomes, and the only difference is static versus fetched.
04THE FIX
Bake the content in
The fix is to make the body synchronously available at build time, without giving up the live fetch for real readers.
- Snapshot the content at build time into a file the component can read synchronously.
- Seed the component's initial state from that snapshot, so the server-rendered HTML already contains the real body.
- Keep the live fetch on mount, so readers still get the freshest version after hydration.
Now the crawler gets a real page, the reader gets live content, and nothing about the browser experience changes.
05VERIFY THE RIGHT LAYER
The serving layer will lie to you
This bug is a specific instance of a trap I keep relearning: you have to know exactly which layer a crawler actually sees, and verify that layer, not the one that is convenient.
It renders is not it is indexed. The browser shows you the best case. Verify the crawler's-eye view, page by page, or the serving layer will lie to you.
06NEXT STEPS
Check your own site
- Curl one content page with scripts off and read the title and body it returns.
- If the body is missing, find where that content is fetched instead of baked.
- Snapshot it at build time and seed the initial render from the snapshot.
- Re-check every page type. Static and fetched bodies behave differently.