September 6, 2026·6 min read

Noindex, Follow: How I Pruned 500 Thin Pages Without Losing Traffic

By Andrew Pyle

Noindex,follow is an HTTP header that tells Google to drop a page from its index while still crawling the links on it. I used it on 514 blog posts that had earned zero clicks in 16 months, and kept the traffic that mattered because the pages that earned it were never touched.

Here is the number that started this. I pulled a 16-month Search Console export for my own site and found that the /blog section was 76% of all clicks. That sounds like a win until you look at what it was ranking for: quantum computing, generic generative-AI filler, Django tutorials. Not one of those is the identity I'm building, which is consulting, SEO, AEO, and GEO. I had a site getting most of its traffic from topics that have nothing to do with what I actually do.

01

573 of 707 posts had never earned a click

I sorted the export by clicks and the shape of the problem got worse. Of 707 blog posts, 573 had zero clicks across the full 16 months. Not low clicks. Zero. Written, published, indexed, sitting in Google's index taking up crawl budget and diluting whatever topical signal the site sends, and not one of them had ever been clicked from search.

That is exactly the shape of thing Google's spam systems target now. The 2026 update demotes thin, near-duplicate, made-for-rankings page mass, and it does not care whether a human or a model wrote the page. It cares whether the page earns its place. Mine mostly didn't. I'm also recovering from a self-inflicted Helpful Content demotion, so a pile of unproven pages sitting in the index isn't neutral. It's radioactive.

02

The decision: suppress 514, keep 133

I reconciled the zero-click set down to 514 posts and noindex,follow'd all of them. The other 133, the ones that had actually earned clicks or meaningful impressions, I left alone. best-free-ai has 23,926 impressions. mcp-certification earns real clicks. Those pages proved something the other 573 never did, so they stay exactly as they were.

This is not a blanket "delete everything under some threshold" move. It's a reconciled list, checked post by post, where the bar is "did this ever earn a click." Anything that cleared the bar even once got a pass. Everything else got suppressed.

03

Noindex, follow, not delete

The mechanism is a single database field: `noindex=True` on the post model. When it's set, nginx emits an `X-Robots-Tag: noindex, follow` header on that page's response, and the page drops out of the sitemap. The URL stays live. Nothing 404s, nothing redirects, nothing gets deleted.

The "follow" half matters as much as the "noindex" half. Noindex alone tells Google to drop the page from search results. Follow tells it to keep crawling the links on that page, so whatever internal link equity the page was passing along keeps flowing instead of dead-ending. It's the difference between pulling a page off the shelf and burning the shelf down.

I checked this the only way that counts: I curled the live site as Googlebot.

``` curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" -I https://andrewjpyle.com/writing/<suppressed-slug> ```

A suppressed page returns 200 with `X-Robots-Tag: noindex, follow` in the response headers. A kept page, like best-free-ai, returns 200 with no such header at all. That's the proof it worked: not "I set a flag and assume it's fine," but the actual header Googlebot sees, verified against the live site.

04

Why not just delete them

Deleting 514 pages felt like the obvious move and I didn't do it, for two reasons.

First, deprecate before delete. If a pruned page turns out to matter, a bug in my reconciliation, a query pattern I didn't account for, a topic that starts earning clicks later, I want an undo button. Clearing `noindex=True` puts the page straight back in the index and the sitemap. Deleting it means starting over from nothing.

Second, some of those 514 pages have internal links pointing at them, and some external links too, however thin. Noindex,follow keeps that link graph intact and lets equity pass through instead of hitting a 404 wall. A page mass problem doesn't require nuking the URLs. It requires the search engine to stop indexing them, which is a narrower and more reversible ask.

This is the same logic I use everywhere else on this site: consolidation runs as redirects instead of deletions, deepening rewrites keep a backup of the original before overwriting it. Reversible by contract, not by good intentions.

05

What this actually buys me

The honest version of the trade: I gave up 0 clicks, because the 514 pruned pages had 0 clicks to give up. What I bought is a cleaner signal. A site where 76% of the indexed pages are actually about the thing the site is for, instead of 76% of the clicks coming from pages about the thing the site isn't for. Google's ranking systems look at a site's pages in aggregate, not just one page in isolation, and 573 zero-click, off-topic pages sitting in the index was dragging on every page next to them.

It also buys me a defensible index. Every page still showing up in Google's results for andrewjpyle.com now has a real, checkable reason to be there: it earned a click, or it's new enough that it hasn't had the chance yet and is being watched. Nothing in the indexed set is there because I published it four years ago and never looked back.

06

This isn't a one-time cleanup, it's a standing rule

The rule going forward is simple and I'm applying it to every page I publish from here: every indexed page earns its place, or it gets noindex,follow'd. Not eventually, not on the next audit. If a page can't show real, per-page substance, a genuine answer to a genuine question, it doesn't get to sit in the index accumulating nothing.

That's the actual bar the 2026 spam update set. It isn't about whether AI touched the page. It's about whether the page is thin, unoriginal, or made for rankings instead of readers. A pruning pass is the retrofit version of that rule, applied to 707 pages I'd already published before I was disciplined about it.

07

The bottom line

Noindex,follow let me pull 514 zero-click pages out of Google's index without deleting a single URL, without breaking a single internal link, and without touching the 133 pages that had actually earned their traffic. The mechanism is boring on purpose: a header, verified live against Googlebot, fully reversible if I got the call wrong. The lesson underneath it is less boring. I'd built a site where three-quarters of the clicks came from topics that weren't mine, and the fix wasn't better content on those pages. It was admitting they didn't belong in the index at all.

Frequently asked questions

What does noindex,follow do?

It tells Google to drop a page from its index while still crawling the links on that page. The page leaves search results while its links keep passing signal.

Why noindex 514 blog posts?

A 16-month Search Console export showed 573 of 707 posts had earned zero clicks, and the blog ranked for topics unrelated to the consulting, SEO, AEO, and GEO identity being built.

Does noindexing pages lose the traffic?

No. The pages that earned traffic were never touched. Only the zero-click thin pages were suppressed, so the traffic that mattered stayed.

Is noindex,follow reversible?

Yes. It is a header change, so a page can be returned to the index later by removing the directive.

Have something you need built or fixed?

I build production Django / Next.js platforms and human-supervised AI-agent systems. Solo, senior, and fast. Tell me what you are building.

Start a project