9,220 pages indexed. Nobody decided to publish them.

Written by

in

A site with 9,220 pages indexed and roughly 8,000 visits a month. Another with 1,630 indexed and a catalogue of a few hundred real products. A third with 1,265 indexable pages, of which 1,082 were blog posts.

None of those businesses set out to publish that much. In every case the platform did it, one URL at a time, over years, and nobody was counting.

That is crawl bloat: the gap between the pages you decided to publish and the pages that exist. It is not a penalty, it is rarely a crawl budget problem, and it is one of the most common things I find.

Where the extra URLs come from

The sources are boringly consistent across platforms.

Expired or discontinued items that keep their page. The car marketplace above was a listings site. Vehicles came and went, which is the business, and the pages stayed, unlinked, in the sitemap, indefinitely.

Filter and facet combinations. Every combination is a URL. The maths is combinatorial, so a store with six filters does not have six extra pages, it has hundreds.

Archives nobody browses. Author archives on a single-author site. Date archives. Tag pages with one post on them. On the fitness site, author archives ran several pages deep for authors with a handful of posts each.

Pagination without canonicals. 266 paginated URLs on that same site, each one a page in Google’s eyes.

URL segments that add nothing. 1,082 posts carrying /blog/ for no reason is not bloat in the count sense, but it is the same inattention showing up in a different column.

Why it matters, stated honestly

Not because Google punishes large sites. It does not.

It matters for a duller reason: everything you do afterwards gets divided by it. Internal links, content effort, whatever authority the domain has: all of it spreads across the pages that exist rather than the pages that matter. On the marketplace, ninety per cent of indexed pages were unreachable, carrying nothing and going nowhere, and every piece of work anyone did on that site was being diluted ten to one.

The second reason is diagnostic. A site where nobody has counted the URLs is a site where nobody has been looking at the technical layer at all, and the rest of the audit usually confirms it.

It also helps to know the base rate before you panic about a number. Ahrefs examined 14 billion pages and found 96.55% get zero traffic from Google. Pages existing and earning nothing is the normal condition of the web, not a symptom peculiar to your site. What makes bloat worth acting on is not that the pages are idle. It is that you are dividing everything you do by them.

What it is not, on most sites, is a crawl budget problem. Google’s own threshold for worrying about that starts at a million pages, or ten thousand changing daily. A 9,220-page site is not close. The budget explanation is available, comforting, and usually wrong.

Counting yours

Three numbers, ten minutes.

What Google has indexed. Search Console, Page indexing, the indexed count.

What you think you have. Ask the person who runs the site. Do not look it up. The discrepancy between belief and reality is itself the finding.

What a crawl finds, with the sitemap fed in as a second source so orphans surface.

If those three numbers are within sight of each other, you do not have this problem. If the first is several times the second, you do.

What I actually recommend

  • Stop the flow before clearing the backlog. This is the recommendation I have changed my mind about. Retro-fixing eight thousand URLs on an old platform is risky and slow. Changing what happens to the next expired listing is one rule in one place, testable before it ships. The backlog stops growing, and in a year the shape of the site is meaningfully better without anyone touching the part nobody understands.
  • Use noindex, not a robots block. Blocking a URL means Google cannot read the noindex on it, so the page stays indexed forever. This is backwards from what most people reach for first.
  • Take them out of the sitemap. A sitemap listing pages you do not want indexed is you actively submitting them.
  • Decide per facet, not per site. Some filtered views are real landing pages people search for. Most are not. That is a judgement per facet and it is the only part of this that takes real thought.

The short version

  1. Crawl bloat is the gap between pages you published and pages that exist.
  2. Expired items, facets, archives and pagination produce almost all of it.
  3. It is not a penalty. It dilutes everything you do afterwards.
  4. It is usually not a crawl budget problem, because Google’s threshold starts at a million pages.
  5. Count three numbers: indexed, believed, crawled. The gaps are the finding.
  6. Stop generating before you clean up. Safer, cheaper, and it holds.
  7. Noindex, not robots.txt. Blocking makes pages permanent.

If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling, indexing and architecture layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

Headshot of Mehul Dedhia, SEO consultant

Written by Mehul Dedhia

I’ve been doing SEO for about fifteen years. Most of my work isn’t building new websites. It’s figuring out why sites that should rank don’t. Local businesses, ecommerce and Shopify brands, affiliate sites, small brands and enterprise. If a site is stuck or has lost rankings, that’s usually where I come in.

Connect with me on LinkedIn