Here are two sites in the same market, from an audit I ran in 2022.
Client
Competitor
Pages indexed
1,510
371
Referring domains
3,575
1,853
Organic traffic / month
23,825
47,722
Keywords in top 3
517
1,216
The client had four times the pages. Nearly twice the referring domains. And half the traffic.
That is not a small gap or a measurement artefact. The competitor was ranking in the top three for more than twice as many keywords off a quarter of the content and half the links.
I’ve thought about this comparison more than any other in fifteen years of audits, because it kills two beliefs at once and most people hold both.
The client wasn’t losing on volume. They were losing on hit rate.
The temptation is to read that table as “the competitor’s content is better” and move on. Look at one more row instead.
Client
Competitor
Keywords ranking in the top 100
26,618
16,142
Keywords ranking in the top 3
517
1,216
The client ranked for ten thousand more keywords than the competitor. They just ranked for almost none of them well.
Roughly 2% of the client’s ranking keywords were in the top three. For the competitor it was about 7.5%. Same market, same query set, and one site converts its rankings into positions that get clicked at nearly four times the rate.
That is the number I’d want on the wall. Not how many keywords you rank for. What proportion of them are in a position anybody sees.
Ranking on page four for twenty-six thousand things is not a foundation to build on. It’s the site telling you it has spread itself across a market rather than winning any part of it.
Where the extra 1,100 pages came from
I went back through the crawl notes for that client, and the extra pages were not a content strategy. They were exhaust.
138 pages under 500 words out of 1,265 indexable
1,082 blog posts carrying /blog/ in the URL for no reason
266 paginated pages with no canonical pointing back to the parent
Author archives, several pages deep, for authors with a handful of posts each
Strip the archives, the pagination and the thin posts and you are not far off the competitor’s page count. The difference between 1,510 and 371 was largely not writing. It was a CMS generating URLs and nobody stopping it.
I’ve written about what that looks like at the extreme — a site where 90% of the indexed pages were unreachable. This is the mild version, and the mild version is far more common.
The link half of it is worse
3,575 referring domains against 1,853, and the smaller profile wins.
Referring domain count is the number everybody quotes because it’s the number every tool puts at the top of the report. It says how many distinct sites link to you. It says nothing whatsoever about whether those sites matter.
Two profiles with identical counts can be completely different assets. One is three thousand five hundred directories, aggregators, scraped listings and expired-domain blogs. The other is eighteen hundred real publications in the sector. The tool renders both as a number, and the bigger number is the weaker profile.
I can’t prove the composition from the data in that audit and I’m not going to pretend otherwise — I didn’t do a link-by-link teardown of a competitor’s profile. But when a site with half your referring domains outranks you this comprehensively, the count is not the variable, and continuing to buy against the count is how the gap gets wider.
The honest counterexample
I don’t want to leave you with “fewer pages and fewer links is better”, because that is just as wrong and I have the data to show it.
Same folder, different audit. A classic car marketplace against the category leader:
Client
Market leader
Pages indexed
9,220
334,000
Referring domains
671
29,812
Organic traffic / month
7,998
497,000
Keywords in top 3
31
11,464
Here the bigger site wins on everything, by a factor of forty or more.
So scale isn’t the enemy. Scale done properly is a genuine moat — 334,000 pages of real inventory, each one a legitimate answer to a specific query, backed by thirty thousand referring domains. Nobody closes that with a blog.
The rule isn’t “smaller is better.” It’s that page count and link count aren’t the variables, and treating them as targets produces sites like the first client — enormous, well-linked, and beaten by something a quarter its size.
What I actually look at instead
When someone shows me a competitor comparison, I ignore the top two rows and go straight to these.
Top-3 keywords as a share of all ranking keywords. The hit rate. It tells you whether the site converts presence into position. Under about 3% and you have a quality or authority problem no amount of publishing fixes.
Traffic per indexed page. Crude, and I use it anyway. The first client was pulling roughly 16 visits per indexed page per month. The competitor was pulling 129. That ratio is the whole story in one number.
How many pages you’d defend. Go through the list and mark every page you would fight to keep. On most sites the honest answer is somewhere under a third. Those are your real page count. Compare that to the competitor’s total instead.
Referring domains you’d name out loud. Same test, for links. How many of those three thousand five hundred would you mention on a sales call?
None of these are clever metrics. They’re the standard numbers with the denominator put back in, and putting the denominator back is most of the work.
What I told that client
Not “write more”. They’d been writing for years and it had produced 26,618 keywords ranking nowhere.
Consolidate the thin posts into fewer, better pages. Fix the URL structure and the pagination so the CMS stops manufacturing URLs. Then put the effort into a smaller number of pages that could realistically reach the top three, and into links from places you’d name on a sales call.
Fewer pages, better placed. It’s a harder sell than a content calendar because it looks like doing less, and for the first three months it is doing less.
The short version
A competitor with 371 pages and 1,853 referring domains beat a client with 1,510 and 3,575 — twice the traffic, more than twice the top-3 rankings.
The client ranked for 10,000 more keywords and almost none of them well. 2% in the top three against 7.5%.
The extra pages weren’t content. Thin posts, pagination, author archives, a bloated URL structure.
Referring domain count says nothing about referring domain quality, and the tools will never tell you the difference.
Scale still wins when it’s real — 334,000 pages of genuine inventory is a moat.
If a site should be ranking and it isn’t, that’s the work I do. If you’re not sure whether your problem is content, links or something underneath both, that’s what an SEO audit is for.
I audited three healthcare clinics this year. Two chiropractic practices and a med spa, in three different towns, with three different web builds and no connection to each other.
Five things were missing from all three.
Three sites is a small sample and I’m not going to pretend otherwise. But these were not five subtle judgement calls. They were five things that a local business is either doing or not doing, and none of the three was doing any of them. That’s fifteen out of fifteen, and after fifteen years I can tell you it’s not a coincidence — it’s what the standard local build leaves out.
Here they are, in the order I’d fix them.
1. The Google Business Profile map is missing, or it’s the wrong map
None of the three had this right.
One had no map embedded anywhere on the site at all. One had it embedded on the home page and the FAQ page, which are not the pages anybody looks for it on, and the embed was pointing at the wrong location. The third had a map on the contact page showing a different business entirely.
Let me be straight about why this is check one, because the usual justification for it is wrong. Embedding a Google map does not make you rank. I’ve seen that claim made a lot and I’ve never seen evidence for it.
What it does is settle the question of which business you are. Your Business Profile, your footer address, your contact page and the map all need to describe one entity in the same terms. When they don’t — and on one of these three they actively contradicted each other — you have handed Google an ambiguity to resolve, and it may not resolve it your way.
It also just helps the customer, who is trying to work out whether you’re close enough to bother with.
Put it in the footer, on the About page, and on the Contact page. Take it from your own Business Profile listing rather than by searching the address, which is how you end up embedding the neighbour.
2. No schema markup
Absent on all three. On one it was completely absent — nothing, anywhere on the site. On another the service pages had none. On the third there was no LocalBusiness markup on the home page.
This is the one I find hardest to be interesting about, because the advice is boring and correct: a local business should have LocalBusiness markup on the home page with name, address, phone, hours, and links out to the profiles that confirm it’s you.
What I will say is what it’s for, since this gets oversold. Schema doesn’t rank you. It tells a machine, unambiguously, facts that it would otherwise have to infer from your page copy. For a local business the facts that matter most — who you are, where you are, when you’re open — are exactly the ones that are usually sitting in a footer image or a styled div where nothing can read them reliably.
Two notes, since both came up in these audits.
Don’t add FAQ schema expecting to see anything. Google restricted FAQ rich results to authoritative government and health sites in 2023, and removed them entirely in May 2026. The Search Console reporting for it is being switched off through this summer. The markup still helps a machine read the page and it does no harm, but if somebody is selling you FAQ schema as a visibility win in 2026, they haven’t read the release notes.
Don’t mark up things that aren’t on the page. Reviews you don’t display, hours you don’t publish. It’s a rich results violation and it’s the fastest way to lose the markup entirely.
3. No reviews page on the website
All three had reviews. One had 44 five-star reviews on its profile, one had 93, one had 8.
None of the three had a page on their own website about them.
I recommend one to every local client, and the reasoning is straightforward. The reviews sitting on your Google Business Profile are on Google’s property. They do their job in the map pack and that’s valuable. But they are not content on your site, they are not something you control, and a person comparing three clinics on a laptop is often reading the websites rather than the profiles.
A reviews page — real quotes, attributed, ideally with the treatment or service named, and a link out to the profile so it can be verified — is a page that can rank, that you own, and that answers the question every prospective patient is actually asking.
The gap between 93 reviews and 8 reviews is also worth naming. The clinic with 8 had a perfectly good practice and no process for asking. That is not an SEO problem and I’d be overreaching to call it one, but it’s the single biggest local visibility difference between those three businesses and no amount of work on the website closes it.
4. No video anywhere
None of the three had a YouTube channel. One of them had a video about their practice on somebody else’s channel, which is a strange position to be in — the content exists and the competitor of nobody in particular owns it.
I want to correct something here that I’ve said myself in the past, including in one of these audits. The usual line is that YouTube is mandatory because Google owns it. That reasoning doesn’t hold up. Ownership isn’t why video shows up.
The real reasons are narrower and they’re enough on their own:
Video appears in search results, in Business Profiles, and increasingly inside AI-generated answers. It’s surface area you don’t otherwise have.
For a treatment nobody has had before, ninety seconds of the practitioner explaining what happens does more for a booking decision than a page of text.
It’s the cheapest trust signal available to a small clinic. A real person, in the real room, on camera.
Three videos is enough to start. What the treatment is, what to expect at a first appointment, who you are.
5. The location question was never answered
All three sites were ambiguous about where they operate. Not wrong — ambiguous. Service pages with no place name anywhere. Contact pages listing four towns. Blog posts targeting one town while the service pages targeted none.
There are two legitimate answers and the sites had chosen neither.
If you want one location, then commit. The service pages carry the town in the URL, the title, the H1 and some of the H2s. Everything on the site points at one place and Google gets a very clear signal.
If you want several, you need a real page for each — with genuinely different content, staff, directions, parking, local specifics. Not a template with the town name swapped out. I’ve seen that build fail more than once, and it fails as a set: forty near-identical pages don’t rank individually and they drag the site down collectively.
The med spa’s contact page named four towns. Nothing else on the site did. That’s the most common version of this — the business knows its catchment perfectly well and the website has never been told.
Pick one. The wrong answer, executed properly, beats no answer.
What wasn’t on the list
Worth saying, because local SEO advice usually leads with these.
Citations weren’t the problem on any of the three. Neither was the primary category — all three were set sensibly. Neither was posting to the Business Profile.
Those get written about constantly and they were the least of it. What was actually missing was more basic and less interesting: tell Google who and where you are, consistently, in every place it looks.
The short version
Five gaps, all three clinics, none of them exotic:
The map embed is missing or pointing at the wrong business. Footer, About, Contact. Take it from your own profile.
No schema. LocalBusiness on the home page. Don’t expect anything visible from FAQ markup in 2026.
No reviews page on your own site. The reviews are Google’s asset until you build one.
No video. Three is enough to start.
No decision about location. One town properly, or real pages for several. Not neither.
None of this is clever. All of it was missing from three out of three practices that were paying for a website and wondering why it wasn’t working.
There is a sixth thing missing from all three, and I left it off the list deliberately because it is a different kind of gap. Not one of them had a page saying which insurers they take or what a visit costs. That does not stop Google understanding the practice, so it never shows up in a crawl — it just quietly costs you the patient who was comparing three clinics and could only find prices on one. I have written about that separately in the sixth gap, and what a clinic should publish about cost.
If a site should be ranking and it isn’t, that’s the work I do. Local SEO is where most of this sits. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for — and if I look and there’s nothing worth fixing, I’ll tell you that.
I went back through eleven audits I’ve written since 2021. Local clinics, a Shopify homeware store, a classic car marketplace, three continuing-education platforms, an ad agency, a car parts retailer.
Different platforms. Different sizes — one had 189 pages, another had over nine thousand. Different budgets, different industries, different countries.
Every one of them had a problem with internal linking. Eleven out of eleven.
Nothing else came close. Schema was missing on ten. Thin content on ten. After that it drops away quickly.
I didn’t expect the number to be that clean, and I’ll be honest that it made me slightly uncomfortable when I counted it. Eleven out of eleven is the kind of statistic that usually means the person counting has defined the thing too loosely.
So I went back and looked at what I’d actually written in each one. They were not the same finding. That’s the part worth your time.
It is never the same problem twice
“Internal linking is weak” is what ends up in the audit document. Underneath it, these were eight genuinely different failures.
1. The money page with no internal links at all.
A med spa’s main service page — the treatment the business is built on — had not one internal link on it. Not to a related service, not back to the home page, not to a booking page. It sat there as a dead end.
That page was expected to rank. Nothing on the site was telling Google it mattered, and nothing on the site gave a visitor anywhere to go once they’d read it.
2. The URL as the anchor text.
Two of the three local sites I audited this year did this throughout their blog. Links written as https://example.com/services/thing/ rather than as words.
The link still passes. It just says nothing. Anchor text is one of the few places you get to tell Google, in plain language, what the destination page is about — and it costs nothing to use. Writing the URL instead is leaving a free field blank.
3. Menu links only, never contextual.
This is the most common version by a distance. The home page links to the service pages, but only through the navigation.
Every page on the site has that navigation. It’s there on the privacy policy. A link from the menu tells Google approximately nothing about the relationship between this page and that one, because the relationship is identical everywhere.
A link from inside the body copy is different. Somebody chose to put it there, in that sentence, with those words around it.
4. Two links to the same page, none to the others.
One clinic’s home page linked to the same service page twice, from two different sections, and to three other service pages not at all.
Nobody decided that. It’s what happens when a page is assembled from blocks over two years by three different people. But the effect is a home page arguing hard for one service and staying silent about the rest.
5. Blog posts that only link to other blog posts.
The blog links to the blog. Post to post to post, a closed circuit.
Meanwhile the pages that actually sell something — the service pages, the collection pages — receive nothing. The blog is doing work and putting it in a cul-de-sac. I found this on both the med spa and one of the clinics.
6. The About page that links only to Contact.
A single internal link on one of the most-visited pages on the site, and it goes to the page a visitor was probably going to find anyway.
7. Orphan pages, at scale.
The extreme end. A classic car marketplace with 8,317 pages that nothing on the site linked to, out of roughly 9,220 indexed. A car parts retailer with 1,277.
At that point you’re not talking about weak internal linking. You’re talking about most of the site being unreachable. I’ve written that one up separately because the cause is different — it’s a database problem, not an editorial one.
8. Too many links.
Worth including because it’s the one people don’t expect. An agency site where every page and post carried a large number of internal and external links in the sidebar.
More links is not better. A page that links to forty things is not emphasising any of them.
Why this one keeps winning
There’s a practical reason internal linking is the finding I write most often, and it isn’t that it’s the most damaging problem. It usually isn’t. A site blocked in robots.txt has a worse day than a site with lazy anchor text.
It’s that internal linking is the only significant ranking input that is entirely inside the building.
Everything else needs somebody else’s cooperation. Links need other people to link to you. Content needs a writer, a brief, a budget and a month. Speed needs a developer. Schema needs somebody who can edit the theme.
Internal links need you to open the page and add a sentence.
That’s the whole argument for calling it the cheapest link building there is. Not that an internal link is worth as much as an external one — it isn’t — but that you can do a hundred of them in an afternoon, for nothing, without asking permission from anybody outside the company.
Which raises the obvious question: if it’s that cheap, why is it broken on eleven sites out of eleven?
Because nobody owns it
In fifteen years I have never once seen internal linking on somebody’s job description.
The developer builds the template. The writer writes the post. The designer decides where the buttons go. The SEO agency is looking at rankings and backlinks. Internal linking is the thing that emerges from all four of them not talking to each other, and it degrades quietly.
It also never breaks visibly. A missing H1 shows up in every crawler. A 404 sends an alert. A page with no internal links looks completely normal. It renders, it loads, it converts if somebody lands on it. There is no red flag anywhere, ever, until somebody sits down and looks.
What I’d actually do, in order
Not a project. An afternoon.
Crawl the site and pull the orphan list first. Sitebulb and Screaming Frog both report this. If the number is large, stop reading this and go find out what’s generating those pages — you have a different problem.
Then find your five most important pages and ask what links to each of them. Not “is it in the menu” — what body-copy link, on what page, with what anchor text. On most sites the answer for at least one of the five is nothing.
Then fix the home page. Contextual links, in the copy, to every service or category that matters, with anchor text that says what the page is. If a service is worth having a page for, it’s worth a sentence on the home page.
Then point the blog at the money pages. Every post should link to at least one page that sells something, in the body, with descriptive anchor text. If a post can’t do that naturally, that tells you something about why the post exists.
Then look at what the menus are promoting. The header and footer are on every page, so whatever is in them is getting the site-wide vote. If the footer has thirty links and four of them matter, the other twenty-six are diluting the four.
Then leave it alone and measure. Internal linking is slow and undramatic. If you change it and rankings move the same week, something else moved.
What it won’t do
I’d rather say this than have you find out in six weeks.
Internal linking will not rescue a site with no authority. If the whole domain has thirty referring domains and the competitors have three thousand, rearranging the links inside your own house is not the constraint. In one of these eleven audits the client had 671 referring domains and the market leader had 29,812. No amount of internal linking closes that.
What internal linking does is make sure the authority you do have lands on the pages you want to rank, rather than being spread evenly across everything you’ve ever published. On most sites that is a meaningful gain and it is available immediately.
It is not a substitute for having something worth linking to.
The short version
It was in every one of eleven audits. Nothing else was.
It’s never the same problem twice — no links at all, URLs as anchors, menu-only, duplicates, blog-to-blog loops, orphans, or simply too many.
It’s the cheapest thing on the list because it needs nobody’s permission.
It’s broken everywhere because nobody owns it, and it never fails visibly.
Orphans first, then the five pages that matter, then the home page, then the blog, then the menus.
It won’t fix an authority problem. It makes the authority you have land where you want it.
If a site should be ranking and it isn’t, that’s the work I do. Link building covers the off-site half. If you’re not sure which of several plausible problems is the one costing you, that’s what an SEO audit is for.
A classic car marketplace had 9,220 pages indexed in Google.
8,317 of them had nothing on the site linking to them.
That’s ninety per cent. Nine pages in every ten existed, were indexed, and could not be reached by clicking. No menu, no category, no related-items block, no link from anywhere. If you didn’t already know the URL, the page was not findable by a person, and it was only findable by Google because the sitemap had handed it over.
The site was live. It was taking traffic — about 8,000 visits a month. Nothing looked wrong. Nobody had noticed, because there is nothing to notice.
What an orphan page actually is
A page nothing links to.
That’s the whole definition, and it’s why the problem hides so well. An orphan page isn’t broken. It returns a 200. It renders. It converts if somebody lands on it. Every tool that checks for errors will tell you it’s fine, because by every measure a tool checks, it is fine.
It’s simply not part of the site any more. It’s a file on a server that happens to be in Google’s index.
The reason that matters isn’t philosophical. Internal links are how importance moves around a website. A page with no internal links receives none of it. It sits at the bottom of the pile regardless of how good it is, and on a large site it’s competing with thousands of siblings in exactly the same position.
They are almost never written by hand
Here’s the thing I’d want you to take from this post, because it changes what you do about it.
Orphan pages are not pages somebody wrote and then forgot to link. That happens, but it’s a handful of pages and it’s not why anybody’s number is in the thousands.
Orphans at scale come out of a database.
The car marketplace was a listings site. Every vehicle got a page. Vehicles came and went — that’s the business — and the pages stayed. Nothing in the template linked to a listing once it dropped off the current inventory. The sitemap kept feeding them to Google. Google kept them.
Nobody made a decision. Nobody wrote 8,317 pages. A system generated them, one at a time, over years.
The same pattern turns up on ecommerce catalogues with discontinued products, job boards with expired postings, event sites, property portals, anything with a feed. If your site has a database behind it, this is your default state unless somebody has designed against it.
The tell: the same number three times
Something in that audit gave the cause away before I’d looked at a single URL.
8,317 pages with no internal links
8,317 pages with a missing canonical, or a canonical still pointing at http
7,353 pages under 500 words
Three findings. Two of them the same number, the third close behind.
That is not three problems. It’s one page-set, generated by one template, carrying every fault of that template. Which is useful, because it means there is one fix, not three — and it means you should stop counting and go and find the template.
I now check for this deliberately. When two numbers in a crawl report match exactly, they are describing the same pages. Chasing them as separate line items is how a technical audit turns into a 400-row spreadsheet that nobody ever implements.
It happens on small sites too
The scale makes the car marketplace memorable, but the proportion is what matters, and the proportion doesn’t need a big site.
Here’s a crawl of a small manufacturer’s ecommerce site — 174 URLs in total.
164 internal HTML pages. 60 orphaned. Thirty-seven per cent, on a site small enough that one person could have clicked every page in an afternoon.
And a car parts retailer, in between the two: 1,630 pages indexed, 1,277 orphans. Seventy-eight per cent. That one also had 1,344 pages under 500 words and no robots.txt file at all, which tells you roughly how much attention the technical layer had been getting.
Three sites, three sizes, same shape.
How to find yours
Two minutes of work, and the reason more people don’t do it is that the check has to be run deliberately — nothing surfaces it on its own.
Crawl the site, then feed the crawler your sitemap as well. This is the step people miss. A crawler starting from your home page follows links. By definition it will never reach a page that has no links pointing at it, so a plain crawl reports zero orphans on a site with eight thousand of them.
You have to give the crawler a second source of URLs — the XML sitemap, a Search Console export, an analytics export, a database dump — and let it compare the two lists. Anything in the second list that the crawl never reached is an orphan.
In Sitebulb it’s under Internal → Orphaned. In Screaming Frog you enable list mode alongside the crawl and use the Orphan URLs report. Both need the sitemap connected first or the number comes back empty and falsely reassuring.
Then sort by whether you care. This is the same question as any other indexing problem: is there anything in that list that was meant to earn something? A thousand expired listings is a maintenance issue. Your third-best service page is an emergency.
What to do with them
There are only three answers, and the first one is the most common.
Let them go. Expired listings, discontinued products, last year’s events. These pages have done their job. Remove them from the sitemap, return a 410 or redirect to the nearest sensible category, and stop generating the URL when the item goes. This is the correct outcome for the large majority of orphans on a database-driven site, and treating it as a loss is why these numbers grow.
Link to them properly. For pages that should exist — a product still on sale, a service page, a location page — the fix is to put them back into the site. Not by adding them to the footer. From the category or hub page they belong under, in the body, with anchor text that describes them. If you can’t work out where a page belongs, that’s an architecture answer, not a linking one.
Consolidate. If forty orphans are forty near-identical thin pages, one good page and forty redirects is better than forty linked thin pages. Linking to something that shouldn’t exist just moves the problem.
What I would not do is add a sitewide “all listings” page with eight thousand links on it. It technically removes the orphan status and it achieves nothing else.
What happened next: nothing
I should be straight about the ending, because it isn’t a success story.
They never did it.
Not because they disagreed with the diagnosis. The site was built on TotalWebManager — a dealer platform old enough and tangled enough that changing how those listing URLs were generated risked breaking things nobody could confidently map. No one could tell me what else depended on that template. The cost of finding out the hard way, on a live site taking real orders, was higher than the cost of leaving 8,317 pages sitting there doing nothing.
I think that was the right call, and I’d rather say so than sulk about it.
An orphan page set is not an emergency. It’s dead weight. It isn’t actively removing you from search the way a stray noindex tag or a robots.txt block does — orphans just quietly fail to help. On a legacy platform where you can’t predict what a template change touches, “leave it alone and don’t make it worse” is a defensible answer. I’ve seen far more damage done by confident fixes to systems nobody understood than by leaving a known problem in place.
What I’d do differently now, on any site like that one: stop the bleeding instead of cleaning up the spill.
Don’t retro-fix eight thousand pages. Change what happens to the next listing when it expires — one rule, one place, testable before it ships. The backlog stays where it is, but it stops growing, and in a year the shape of the site is meaningfully better without anybody having touched the risky part.
That’s the version of this recommendation I now give whenever the platform is old enough that nobody wants to be the person who broke it.
The honest caveat
Removing orphan pages is a cleanup, not a growth lever. I have never seen a site’s traffic jump because orphans were dealt with in isolation, and I would be suspicious of anyone who told you otherwise.
What it does is stop you wasting effort. On the car marketplace, ninety per cent of the indexed pages were carrying nothing and going nowhere, and every piece of work anybody did on that site — content, links, speed — was being spread across ten times more pages than needed to exist. The value of fixing it is that everything you do afterwards lands somewhere.
It’s also the fastest way to find out that your real problem is architectural. A site with a 90% orphan rate does not have an internal linking problem. It has a site that was never designed to have a shape.
The short version
An orphan page is a page nothing links to. It isn’t broken, which is why nothing flags it.
At scale they come from a database, not a writer. Listings, expired products, feeds.
Matching numbers in a crawl report mean one page-set, not several problems. 8,317 orphans and 8,317 bad canonicals were the same pages.
A plain crawl will report zero. You must give the crawler the sitemap as a second URL source.
Ask what you actually care about before you touch anything.
Most of them should be removed, not linked. Some should be linked properly. Some should be consolidated.
It’s a cleanup, not a growth lever — but it’s what makes the growth work land.
On an old platform, fixing the flow beats fixing the backlog. Stop generating new orphans; leave the existing ones if touching them is risky.
If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling, indexing and architecture layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.
A med spa came to me because the site wasn’t showing up locally. Good clinic, real reviews, services people search for every day.
The address in their footer linked to a Google Maps listing for a different business. So did the address on the contact page. The map embedded on the contact page was that other business too, and the Instagram icon in the footer went to that company’s account.
Nobody had done anything wrong on purpose. A previous build had been adapted from another clinic’s site and those links came with it. But for anyone arriving — a customer or a crawler — the site’s own contact details pointed somewhere else.
That took about four minutes to find, and no keyword tool was involved.
This is what I do when someone sends me a URL and says the site should be ranking and isn’t. I start on the home page and work down. The whole first pass takes up to three days, and almost none of it is spent where people expect.
Here’s the order, and more usefully, why it’s in that order.
Before anything: start the crawl
The first thing isn’t really a check. I put the URL into Sitebulb — Screaming Frog does the same job — and let it run.
A crawl on a real site takes time, and there’s no sense watching it. So it churns while I look at the site the way a customer would, and I come back once there’s something to read.
Worth saying, because published checklists always present themselves as a clean sequence and real work isn’t one. The crawl is the fourth thing I read. It’s the first thing I start.
1. The home page
The most important page on the site, and where I always begin.
I think of the home page as a movie trailer. It isn’t the film. Its job is to show enough of what the business is — snippets, glimpses, a sense of the thing — that a stranger decides to stay. A real photo of the owner. Testimonials from people who exist. Enough of the story that the business reads as legitimate.
It’s the first impression when you meet someone. You know within seconds whether you’re going to keep talking.
Does it say what the business actually is? Plainly, near the top, in words a customer would use. A great many home pages describe a mood rather than a business.
Does it link down to the pages that matter? This is the one I find broken most often. On a local site the home page should link contextually to the service pages, inside the copy, not only from the menu. On a Shopify store, to the collection pages.
Menu links are fine, but every page has those. A contextual link from the home page body, with anchor text that describes the destination, says something different about what that page is for.
Two variants of this I’ve found recently. One clinic’s home page linked to the same service page twice, from two different places, and to three other services not at all. Another home page had plenty of text but none of it mentioned a single service by name.
Is the structure sane? A chiropractic clinic I looked at had seven H1 tags on the home page, several of them empty. There were blank H2 tags too. That’s not somebody’s SEO strategy, it’s a page builder emitting headings for layout reasons.
I mention this because it’s invisible. The page looked fine. It looked good, in fact. You only see it if you look at the markup, which is why it survives years of people wondering why the site doesn’t rank.
Is there a testimonial section and an FAQ? Both earn a place.
I’ll be straight about why this is check one rather than something more technical: a home page that doesn’t establish trust makes everything downstream a waste of effort. Rank a page like that better and you’ve only sent more people somewhere to bounce off. That’s a conversion problem more than a ranking one — but it’s the reason I won’t spend three days on a site and skip it.
2. The header and footer menus
Then the navigation, on the assumption nobody has looked at it since the theme was installed.
I check that links work, where they point, and what’s being given prominence.
On one clinic site, a header menu item labelled Athlete Care pointed at /elementor-685/. That’s an unfinished page-builder draft. It was in the main navigation of a live site, on every page, being crawled every time.
On a Shopify store I audited, the header menu item names had all been assigned H2 tags by the theme. So had every product name on the home page and collection pages. A collection page with forty products was handing Google forty H2 headings that were just product names, before any of the actual page content.
That’s a theme-level default, it’s on thousands of stores, and almost nobody looks for it.
Elsewhere, a clinic’s blog posts didn’t display the header menu at all. Every post on that site was a dead end — a visitor landing there from search had no way through to a service page except the back button.
Home, about, contact, blog and the real target pages should be reachable from the menus. Often one of them isn’t.
3. Reading the crawl
By now the crawl has finished. This is where the mechanical problems surface — broken pages, multiple or missing H1s, duplicate titles and descriptions, title lengths, canonicals, noindexed URLs, external links, orphan pages.
Most of that is ordinary and every crawler reports it. Three things are worth more attention than they get.
Scale. The numbers on a real site are not what people imagine. One Shopify store came back with 1,562 URLs with a missing or empty H1. An auto accessories store had 1,301 URLs missing an H1, and 1,347 pages with either no canonical at all or a canonical still pointing at the http version. The same 1,347 were missing a meta description.
Nobody creates 1,562 problems by hand. At that scale it’s always one template or one setting, and finding what produced it matters more than the list.
Pages the platform invented. That Shopify store had 659 tag pages, competing with the collection pages that were meant to rank. Nobody wrote them. Nobody linked to them deliberately. They exist because the platform makes them.
Orphan pages. These are pages nothing links to — not from the menu, not from another page, not from anywhere. A visitor can only reach one by typing the URL.
The auto accessories store had 1,277 orphan pages against 1,630 pages indexed in Google. Most of that site was unreachable from the rest of it.
That’s the finding I most often have to explain twice, because every one of those pages exists, loads, and looks completely normal when you open it.
And it isn’t only a big-site problem. Here’s the internal URL summary from a crawl of a small outdoor gear store:
One hundred and seventy-four URLs in total. Sixty of them orphaned. Sixty-eight redirects and sixteen broken pages on a site you could read end to end in an afternoon.
Over a third of that store was pages nothing pointed at. The owner knew about none of it, and no amount of writing better product descriptions would have touched it.
4. Redirects
Short, and occasionally it answers the whole question.
The www version should resolve to non-www, or the other way round — I don’t much mind which, as long as it’s consistent. HTTP should redirect to HTTPS. Canonicals should be self-referencing and on the https version, which as above is frequently not the case.
When this is wrong it’s usually been wrong for years, and it means the site has quietly been running as more than one site.
5. Is there enough site here at all?
Then I stand back and ask whether there’s enough substance to rank on.
This is a judgement, and it differs for local, ecommerce and national sites. But I carry rough tripwires:
fewer than about 50 well-written blog posts and the site probably hasn’t said enough to be taken seriously on its topic
a service or collection page at 300 words with a single heading almost certainly hasn’t covered what a buyer needs. The ones that work tend to run longer, with an H1, several H2s and some H3s underneath
target pages generally want an FAQ section
A medical education site I audited had 189 indexable pages, of which 93 were under 500 words. Half the site was pages that didn’t say enough to be worth indexing.
Now the important part. These are not targets. There is no ideal page length, Google has said so plainly, and writing to a word count produces exactly the padded, say-nothing page that got the site here.
They’re tripwires. When a page falls well under them, it’s a reliable sign nobody has properly answered the question that page exists to answer. The number tells me where to look. Reading it tells me whether there’s a problem.
Pad a thin page to 1,200 words and you have a longer thin page.
6. Internal linking
I open a handful of pages at random and read the links inside the content.
What I find, over and over: pages with no contextual internal links at all, pages with far too many, generic anchor text — “click here”, “read more” — and naked URLs used as anchors.
Two examples from one med spa. A hormone therapy service page had no internal links whatsoever. A weight loss service page had exactly one, to the contact page. The blog posts linked to each other using the raw URL as the anchor text, and never linked to a service page at all.
That’s a site publishing content that supports nothing.
Internal linking is the part of SEO nobody needs permission or budget for. It costs nothing, it’s entirely within the owner’s control, and it’s neglected on most of the sites I’m called into.
7. The about page
Almost always thin, and almost always disconnected.
Two checks: does it link back to the home page, and to contact. Usually neither — one med spa’s about page had a single internal link on it, to the contact page, and nothing else. Another site didn’t have an about page at all.
The bigger problem is what’s on it. Most about pages say the business was established in a year and is committed to quality. That isn’t a story and it tells nobody anything.
Every brand has a story. Why the person started, what they did before, who the team are. That’s where the trust the home page promised gets substantiated or doesn’t.
8. The contact page
Same pattern. Thin, and treated as a form rather than a page.
I look for a link back to the home page, the Google Business Profile map embedded, and the address and phone number present as text.
On local sites this is where the damage tends to be. Beyond the med spa at the top of this post, the recurring one is subtler: the address on the site doesn’t exactly match the address on the Google Business Profile. Not wrong — a suite number formatted differently, an abbreviation, a missing line. It’s the sort of difference a human wouldn’t notice and a machine can’t ignore.
Two of the local sites I looked at recently had no GBP map embedded anywhere at all. One had the map embedded, incorrectly, in two places.
9. Reviews and the profile
On local work I look at the Google Business Profile alongside the site, because ranking in the map pack and ranking the website are two different jobs and clients rarely separate them.
Mostly I’m looking at review volume against the competition rather than in isolation. One clinic had 93 five-star reviews, which sounds excellent until you look at who they’re competing with. A med spa had 8. Neither number means anything until you’ve looked at the other businesses in the pack.
What I don’t look at for the first three days
You’ll have noticed there’s nothing here about backlinks.
That’s deliberate. On-page is the foundation of a website, and I want to know the foundation is sound before I look at anything built on top of it. Links pointing at a site whose home page doesn’t say what the business does, whose service pages have no internal links, and a third of whose URLs are orphaned are not going to fix any of that. They’ll send authority into a structure that can’t carry it.
I do get to off-page work. Comparing a site against the three or four businesses beating it — traffic, ranking keywords, referring domains, pages indexed — is a real part of a full audit and sometimes it’s where the actual answer is. Occasionally a site is doing everything right on-page and is simply outgunned, and the honest finding is that it needs links and time rather than another round of fixes. But I can’t tell that from the outside until I know the site itself is coherent, because otherwise I’m attributing to competition what’s really a broken template.
So it’s sequencing, not dismissal. Foundation, then everything else.
The second thing missing is Search Console, and that one’s practical rather than philosophical.
A proper Search Console audit tells you things nothing else can — what the site is already getting impressions for, which pages Google has actually indexed, which queries sit in positions four to twenty and need a nudge rather than a rewrite. It’s some of the most useful data available on any site.
But it’s the client’s data, and I need them to grant me access before I can see any of it.
Which is why the three days above are structured the way they are. Everything on this list can be done from the outside, with nothing but a URL. No logins, no access requests, no waiting on an agency to release something. That’s the entire point of the first pass — it tells me whether there’s a real problem worth either of us spending money on, before anyone has signed anything.
What happened
Two of these, a Shopify store and a local clinic.
Same pattern in both. Nothing moved for three to four weeks after the work went in, then it started. On the local site the service pages began showing up. On the Shopify store the collection pages and the home page both improved in rankings and impressions.
The three-to-four week gap matters, because that’s the point where most people conclude the work didn’t help. It isn’t long enough to draw a conclusion from. It’s roughly how long it takes for a site to be recrawled, reassessed, and for anything to surface.
Why I’m not going to claim one change caused that
Here’s where I part company with how these stories usually get told.
The convention is a before-and-after graph, a green arrow, and an implied straight line from one change to the recovery. I’m not going to give you that, because I don’t believe it and neither should you.
Rankings move for many reasons at once. Sites get worked on while other things happen to them — an update lands, a competitor changes, seasonality does what seasonality does, the client runs a campaign nobody mentioned. When a set of fixes goes in together and things improve a month later, the honest description is that things improved after the work, not because of any single item in it.
Anyone showing you a graph proving one change caused a recovery is guessing with more confidence than the data supports.
What I can tell you is narrower and more useful: these are the problems that were there, this is the order I found them in, and this is roughly how long before anything moved.
The short version
Up to three days, in this order:
Start the crawl and leave it running.
The home page — what the business is, whether it earns trust, whether it links contextually to service or collection pages, and how many H1s it’s actually emitting.
Header and footer menus — working, pointing somewhere real, and not handing out heading tags for layout.
Read the crawl — missing H1s, duplicate and missing titles and descriptions, canonicals, noindexed URLs, platform-generated pages, orphans.
Redirects — one canonical host, HTTPS enforced, self-referencing canonicals.
Is there enough site here — judged by reading, not counting.
Internal linking — contextual links inside the content, anchor text that means something.
About page — linked, substantial, an actual story.
Here is the Page Indexing report for a Shopify store I have worked on.
Four hundred and seventy-nine pages indexed. Six and a half thousand not.
That is not a broken site. It sells, it ranks, it has a real business behind it. It is a normal Shopify store, and the ratio is normal too — which is the part worth sitting with before you do anything about it.
Open the same report on your own site and there is a row called Crawled — currently not indexed. On most sites it is not a small number. On a Shopify store it is routinely larger than the number of pages you think you have.
The status itself is unambiguous. Google found the URL, spent the resource to crawl it, looked at what came back, and decided not to put it in the index. There was no error. Nothing is broken. Google simply did not think the page was worth keeping.
What happens next is almost always the same. The client sees a number in the thousands, reads it as a verdict on the site, and wants it fixed.
In my experience that number is mostly pages the platform created on its own. Nobody wrote them. Nobody linked to them. Nobody would miss them. Google has looked at them and declined, which is the correct outcome, and the report is doing its job.
So before any of the diagnosis below, there is one question that settles most of these conversations: is there anything in that list you actually care about?
First, find out whether any of it matters
Do not start at the top of the list and work down. Start by looking for pages that were meant to earn something:
the home page
service pages
collection pages
product pages
location pages
blog posts
If none of those are in there, you are looking at platform exhaust. The work is a cleanup job — stop generating the pages — and not an investigation. A great many clients have spent months worrying about a number that meant nothing was wrong.
If one of those pages is in the list, open it yourself. On a laptop and on a phone.
Not in a crawler. Not in a tool. In a browser, on both, as a customer would meet it. Mobile matters more than laptop here, because that is predominantly what Google is indexing from — a page that renders fine on a desktop and collapses on a phone is not a mysterious indexing problem. It is a broken page, and you have just found the cause in under a minute.
A surprising number of these are resolved at exactly that point, because the page is visibly not what the person believed they had published.
If the page loads properly on both, submit it for indexing in Search Console and watch what happens.
Then ask whether the page should exist
Here is the breakdown for that same store.
1,241 pages crawled and not indexed. That alone is more than twice the number of pages Google has actually indexed.
But look at the row above it. 3,292 excluded by a noindex tag — half the entire report. Somebody, at some point, told Google not to index those. Almost certainly not deliberately, one by one. That is the theme doing it, or an app, or a setting nobody has revisited.
Most of the URLs in these reports were never deliberately created. They are a by-product of the platform.
Shopify is the clearest example, because it generates a great deal on your behalf. Collection URLs that duplicate each other. Product URLs available under several paths. Tag pages. Filter combinations. Vendor pages nobody linked to. None of it was a decision anyone made — it is just what the platform does when you add products.
WordPress does the same thing more quietly. Author archives on a single-author site. Date archives nobody will ever browse. Tag pages with one post in them. Attachment pages, if nothing has been done about them.
So the first question is not how do I improve this page. It is did I mean to publish this page at all.
If the answer is no, there is nothing to improve. Google has already made the correct decision about it, and you are looking at a report that is functioning exactly as intended. The work is to stop generating the pages, not to fix them.
That reframing removes a large share of the report on most sites before you have diagnosed anything.
When it is a content problem, improve it or delete it
For the pages you did mean to publish, thin content is the main cause. I do not have a contrarian position on this. The standard advice is correct.
Where I differ is on the second option. The advice is usually given as improve the content, and improvement is treated as the only acceptable answer. It is not. For a lot of pages the honest answer is that they should not have been written, and the correct action is to remove them.
A page that exists because someone decided the site needed more pages is not going to become valuable by having more words added to it. Google has looked at it and told you what it thinks. Adding four hundred words does not change the underlying fact that the page has nothing to say.
Improve the ones with a real reason to exist. Delete the rest. Both are legitimate outcomes, and treating deletion as failure is why these reports grow.
The causes that are not about content at all
The remainder are technical, and they are the ones that get missed because everyone has already accepted the content explanation.
The sitemap was never submitted. It happens far more often than it should, including on sites that have had an agency for years. It does not directly cause this status, but on a large site with weak internal linking it changes what Google prioritises, and it is a thirty-second check.
Googlebot is spending its time on the wrong files. In Search Console, under Settings, the Crawl Stats report breaks requests down by file type. Here is a WordPress site — a local business, fewer than a hundred pages worth indexing:
JavaScript 34%. CSS 26%. HTML 26%.
Googlebot spent more of its time on this site fetching stylesheets and scripts than fetching pages. Only about a quarter of everything it requested was a page at all. Google is not refusing to index the content so much as never getting a clear run at it, because most of what it collected was furniture.
That is a plugin and theme problem, not a writing problem, and no amount of rewriting will touch it. It is also two clicks from the report everyone is already staring at, and almost nobody looks at it.
A warning about reading that chart. If you are looking at a domain property rather than a single host, the breakdown covers every subdomain together. I have a client property showing JavaScript at 95% and HTML at 1%, which looks catastrophic until you open the host list and find that a static asset subdomain accounts for 1.1 million of the 1.16 million requests. The main site is fine. The chart was describing a CDN.
So segment by host before you conclude anything. The file-type chart is one of the more useful screens in Search Console and one of the easiest to misread.
The server is not answering reliably. The same screen has a Host Status line covering robots.txt fetching, DNS resolution and server connectivity.
Look at what that one says: host had problems in the past. It is a green tick with a caveat attached, and it is very easy to scroll past a green tick.
Intermittent failures do not always surface as errors anywhere else, and they change how much Google is willing to attempt. Worth opening rather than accepting the tick.
How long it takes once you have fixed it
Around three weeks, in most cases, for pages to start being indexed after the actual cause has been addressed.
Not three weeks from when you make the change. Three weeks from when you make the right change, which is a different date and usually a later one.
If it has been six weeks and nothing has moved, the assumption to revisit is not the timeline. It is the diagnosis.
The short version
The status is not a verdict on your writing. It is a description of what Google did.
Work through it in this order:
Is anything in the list important? Home page, service, collection, product, location page, blog post. If none of those are there, stop. Nothing is wrong.
If one of them is there, open it on a laptop and a phone. Most problems are visible at this point, and mobile is where they show up.
If it loads properly on both, submit it for indexing and watch what happens.
If the page is one you meant to publish and it is thin, improve it or delete it. Both are fine.
If it is none of those, go technical. Sitemap, crawl distribution by file type, host status.
Give it three weeks from the correct fix, then re-examine the diagnosis rather than waiting longer.
Most of the sites I am called into have spent months on step four, for pages that never should have got past step one.
If a site should be ranking and is not, that is the work I do. Technical SEO covers the crawling, rendering and indexing layer. If you are not sure which of several plausible problems is the one costing you, that is what an SEO audit is for.