Author: Mehul Dedhia

  • Eleven audits, fifteen years, one finding in every single one

    I went back through eleven audits I’ve written since 2021. Local clinics, a Shopify homeware store, a classic car marketplace, three continuing-education platforms, an ad agency, a car parts retailer.

    Different platforms. Different sizes. One had 189 pages, another had over nine thousand. Different budgets, different industries, different countries.

    Every one of them had a problem with internal linking. Eleven out of eleven.

    Nothing else came close. Schema was missing on ten. Thin content on ten. After that it drops away quickly.

    I didn’t expect the number to be that clean, and I’ll be honest that it made me slightly uncomfortable when I counted it. Eleven out of eleven is the kind of statistic that usually means the person counting has defined the thing too loosely.

    So I went back and looked at what I’d actually written in each one. They were not the same finding. That’s the part worth your time.

    It is never the same problem twice

    “Internal linking is weak” is what ends up in the audit document. Underneath it, these were eight genuinely different failures.

    1. The money page with no internal links at all.
    A med spa’s main service page, the treatment the business is built on, had not one internal link on it. Not to a related service, not back to the home page, not to a booking page. It sat there as a dead end.

    That page was expected to rank. Nothing on the site was telling Google it mattered, and nothing on the site gave a visitor anywhere to go once they’d read it.

    2. The URL as the anchor text.
    Two of the three local sites I audited this year did this throughout their blog. Links written as https://example.com/services/thing/ rather than as words.

    The link still passes. It just says nothing. Anchor text is one of the few places you get to tell Google, in plain language, what the destination page is about, and it costs nothing to use. Writing the URL instead is leaving a free field blank.

    The automated checks will not save you here. HTTP Archive’s 2025 Web Almanac found 86.6% of mobile home pages pass Lighthouse’s descriptive-link-text audit, which sounds like a solved problem. It is not. That audit catches “click here” and “read more”. It does not flag a raw URL used as anchor text, and it has nothing at all to say about a page whose only links are the ones in the menu. Nine sites in ten pass a test that was never measuring the thing that is wrong with them.

    3. Menu links only, never contextual.
    This is the most common version by a distance. The home page links to the service pages, but only through the navigation.

    Every page on the site has that navigation. It’s there on the privacy policy. A link from the menu tells Google approximately nothing about the relationship between this page and that one, because the relationship is identical everywhere.

    A link from inside the body copy is different. Somebody chose to put it there, in that sentence, with those words around it.

    There is a pattern in the aggregate data worth knowing about. In the 2024 Web Almanac, the median page across the whole web carried 41 internal links. On the top thousand sites by traffic, the median was 129. Three times as many.

    I would not read that as cause and effect — large sites have more pages to link to, and they got large for reasons that have nothing to do with linking. But hold it next to the eleven audits and the shape is the same one. The sites at the top are doing a great deal of this. The sites I get called into are doing almost none of it, and nothing except attention is stopping them.

    4. Two links to the same page, none to the others.
    One clinic’s home page linked to the same service page twice, from two different sections, and to three other service pages not at all.

    Nobody decided that. It’s what happens when a page is assembled from blocks over two years by three different people. But the effect is a home page arguing hard for one service and staying silent about the rest.

    5. Blog posts that only link to other blog posts.
    The blog links to the blog. Post to post to post, a closed circuit.

    Meanwhile the pages that actually sell something, the service pages and the collection pages, receive nothing. The blog is doing work and putting it in a cul-de-sac. I found this on both the med spa and one of the clinics.

    6. The About page that links only to Contact.
    A single internal link on one of the most-visited pages on the site, and it goes to the page a visitor was probably going to find anyway.

    7. Orphan pages, at scale.
    The extreme end. A classic car marketplace with 8,317 pages that nothing on the site linked to, out of roughly 9,220 indexed. A car parts retailer with 1,277.

    At that point you’re not talking about weak internal linking. You’re talking about most of the site being unreachable. I’ve written that one up separately because the cause is different. It’s a database problem, not an editorial one.

    8. Too many links.
    Worth including because it’s the one people don’t expect. An agency site where every page and post carried a large number of internal and external links in the sidebar.

    More links is not better. A page that links to forty things is not emphasising any of them.

    Why this one keeps winning

    There’s a practical reason internal linking is the finding I write most often, and it isn’t that it’s the most damaging problem. It usually isn’t. A site blocked in robots.txt has a worse day than a site with lazy anchor text.

    It’s that internal linking is the only significant ranking input that is entirely inside the building.

    Everything else needs somebody else’s cooperation. Links need other people to link to you. Content needs a writer, a brief, a budget and a month. Speed needs a developer. Schema needs somebody who can edit the theme.

    Internal links need you to open the page and add a sentence.

    That’s the whole argument for calling it the cheapest link building there is. Not that an internal link is worth as much as an external one, because it isn’t, but that you can do a hundred of them in an afternoon, for nothing, without asking permission from anybody outside the company.

    Which raises the obvious question: if it’s that cheap, why is it broken on eleven sites out of eleven?

    Because nobody owns it

    In fifteen years I have never once seen internal linking on somebody’s job description.

    The developer builds the template. The writer writes the post. The designer decides where the buttons go. The SEO agency is looking at rankings and backlinks. Internal linking is the thing that emerges from all four of them not talking to each other, and it degrades quietly.

    It also never breaks visibly. A missing H1 shows up in every crawler. A 404 sends an alert. A page with no internal links looks completely normal. It renders, it loads, it converts if somebody lands on it. There is no red flag anywhere, ever, until somebody sits down and looks.

    What I’d actually do, in order

    Not a project. An afternoon.

    1. Crawl the site and pull the orphan list first. Sitebulb and Screaming Frog both report this. If the number is large, stop reading this and go find out what’s generating those pages. You have a different problem.
    2. Then find your five most important pages and ask what links to each of them. Not “is it in the menu”, but what body-copy link, on what page, with what anchor text. On most sites the answer for at least one of the five is nothing.
    3. Then fix the home page. Contextual links, in the copy, to every service or category that matters, with anchor text that says what the page is. If a service is worth having a page for, it’s worth a sentence on the home page.
    4. Then point the blog at the money pages. Every post should link to at least one page that sells something, in the body, with descriptive anchor text. If a post can’t do that naturally, that tells you something about why the post exists.
    5. Then look at what the menus are promoting. The header and footer are on every page, so whatever is in them is getting the site-wide vote. If the footer has thirty links and four of them matter, the other twenty-six are diluting the four.
    6. Then leave it alone and measure. Internal linking is slow and undramatic. If you change it and rankings move the same week, something else moved.

    What it won’t do

    I’d rather say this than have you find out in six weeks.

    Internal linking will not rescue a site with no authority. If the whole domain has thirty referring domains and the competitors have three thousand, rearranging the links inside your own house is not the constraint. In one of these eleven audits the client had 671 referring domains and the market leader had 29,812. No amount of internal linking closes that.

    What internal linking does is make sure the authority you do have lands on the pages you want to rank, rather than being spread evenly across everything you’ve ever published. On most sites that is a meaningful gain and it is available immediately.

    It is not a substitute for having something worth linking to.

    The short version

    1. It was in every one of eleven audits. Nothing else was.
    2. It’s never the same problem twice: no links at all, URLs as anchors, menu-only, duplicates, blog-to-blog loops, orphans, or simply too many.
    3. It’s the cheapest thing on the list because it needs nobody’s permission.
    4. It’s broken everywhere because nobody owns it, and it never fails visibly.
    5. Orphans first, then the five pages that matter, then the home page, then the blog, then the menus.
    6. It won’t fix an authority problem. It makes the authority you have land where you want it.

    If a site should be ranking and it isn’t, that’s the work I do. Doing this properly, across a whole site, is contextual internal linking. Link building covers the off-site half. If you’re not sure which of several plausible problems is the one costing you, that’s what an SEO audit is for.

  • Ninety per cent of this site was unreachable

    A classic car marketplace had 9,220 pages indexed in Google.

    8,317 of them had nothing on the site linking to them.

    That’s ninety per cent. Nine pages in every ten existed, were indexed, and could not be reached by clicking. No menu, no category, no related-items block, no link from anywhere. If you didn’t already know the URL, the page was not findable by a person, and it was only findable by Google because the sitemap had handed it over.

    The site was live. It was taking traffic, about 8,000 visits a month. Nothing looked wrong. Nobody had noticed, because there is nothing to notice.

    What an orphan page actually is

    A page nothing links to.

    That’s the whole definition, and it’s why the problem hides so well. An orphan page isn’t broken. It returns a 200. It renders. It converts if somebody lands on it. Every tool that checks for errors will tell you it’s fine, because by every measure a tool checks, it is fine.

    It’s simply not part of the site any more. It’s a file on a server that happens to be in Google’s index.

    The reason that matters isn’t philosophical. Internal links are how importance moves around a website. A page with no internal links receives none of it. It sits at the bottom of the pile regardless of how good it is, and on a large site it’s competing with thousands of siblings in exactly the same position.

    They are almost never written by hand

    Here’s the thing I’d want you to take from this post, because it changes what you do about it.

    Orphan pages are not pages somebody wrote and then forgot to link. That happens, but it’s a handful of pages and it’s not why anybody’s number is in the thousands.

    Orphans at scale come out of a database.

    The car marketplace was a listings site. Every vehicle got a page. Vehicles came and went, which is the business, and the pages stayed. Nothing in the template linked to a listing once it dropped off the current inventory. The sitemap kept feeding them to Google. Google kept them.

    Nobody made a decision. Nobody wrote 8,317 pages. A system generated them, one at a time, over years.

    The same pattern turns up on ecommerce catalogues with discontinued products, job boards with expired postings, event sites, property portals, anything with a feed. If your site has a database behind it, this is your default state unless somebody has designed against it.

    The tell: the same number three times

    Something in that audit gave the cause away before I’d looked at a single URL.

    • 8,317 pages with no internal links
    • 8,317 pages with a missing canonical, or a canonical still pointing at http
    • 7,353 pages under 500 words

    Three findings. Two of them the same number, the third close behind.

    That is not three problems. It’s one page-set, generated by one template, carrying every fault of that template. Which is useful, because it means there is one fix, not three, and it means you should stop counting and go and find the template.

    I now check for this deliberately. When two numbers in a crawl report match exactly, they are describing the same pages. Chasing them as separate line items is how a technical audit turns into a 400-row spreadsheet that nobody ever implements.

    The canonical half of that is not exotic either. HTTP Archive’s 2025 Web Almanac found a canonical tag on only 67% of mobile pages — a third of the web has none. What made the car marketplace unusual was not that a template omitted the tag. It was that one template was responsible for nine tenths of the site, so a single omission printed itself 8,317 times.

    It happens on small sites too

    The scale makes the car marketplace memorable, but the proportion is what matters, and the proportion doesn’t need a big site.

    Here’s a crawl of a small manufacturer’s ecommerce site, 174 URLs in total.

    Sitebulb crawl summary showing 164 internal HTML pages, of which 60 are orphaned, alongside 16 broken URLs and 68 redirects

    164 internal HTML pages. 60 orphaned. Thirty-seven per cent, on a site small enough that one person could have clicked every page in an afternoon.

    And a car parts retailer, in between the two: 1,630 pages indexed, 1,277 orphans. Seventy-eight per cent. That one also had 1,344 pages under 500 words and no robots.txt file at all, which tells you roughly how much attention the technical layer had been getting.

    Three sites, three sizes, same shape.

    On the robots.txt point, before anyone assumes that was freakish: the Web Almanac found 13% of sites return a 404 for robots.txt. One in eight. Which is not itself a crisis — no robots.txt means no restrictions, and plenty of small sites need none. The tell is not the missing file. It is a 1,630-page ecommerce catalogue where nobody had ever had cause to open one.

    How to find yours

    Two minutes of work, and the reason more people don’t do it is that the check has to be run deliberately. Nothing surfaces it on its own.

    Crawl the site, then feed the crawler your sitemap as well. This is the step people miss. A crawler starting from your home page follows links. By definition it will never reach a page that has no links pointing at it, so a plain crawl reports zero orphans on a site with eight thousand of them.

    You have to give the crawler a second source of URLs: the XML sitemap, a Search Console export, an analytics export, a database dump. Then let it compare the two lists. Anything in the second list that the crawl never reached is an orphan.

    In Sitebulb it’s under Internal → Orphaned. In Screaming Frog you enable list mode alongside the crawl and use the Orphan URLs report. Both need the sitemap connected first or the number comes back empty and falsely reassuring.

    Then sort by whether you care. This is the same question as any other indexing problem: is there anything in that list that was meant to earn something? A thousand expired listings is a maintenance issue. Your third-best service page is an emergency.

    What to do with them

    There are only three answers, and the first one is the most common.

    • Let them go. Expired listings, discontinued products, last year’s events. These pages have done their job. Remove them from the sitemap, return a 410 or redirect to the nearest sensible category, and stop generating the URL when the item goes. This is the correct outcome for the large majority of orphans on a database-driven site, and treating it as a loss is why these numbers grow.
    • Link to them properly. For pages that should exist, a product still on sale, a service page, a location page, the fix is to put them back into the site. Not by adding them to the footer. From the category or hub page they belong under, in the body, with anchor text that describes them. If you can’t work out where a page belongs, that’s an architecture answer, not a linking one.
    • Consolidate. If forty orphans are forty near-identical thin pages, one good page and forty redirects is better than forty linked thin pages. Linking to something that shouldn’t exist just moves the problem.

    What I would not do is add a sitewide “all listings” page with eight thousand links on it. It technically removes the orphan status and it achieves nothing else.

    What happened next: nothing

    I should be straight about the ending, because it isn’t a success story.

    They never did it.

    Not because they disagreed with the diagnosis. The site was built on TotalWebManager, a dealer platform old enough and tangled enough that changing how those listing URLs were generated risked breaking things nobody could confidently map. No one could tell me what else depended on that template. The cost of finding out the hard way, on a live site taking real orders, was higher than the cost of leaving 8,317 pages sitting there doing nothing.

    I think that was the right call, and I’d rather say so than sulk about it.

    An orphan page set is not an emergency. It’s dead weight. It isn’t actively removing you from search the way a stray noindex tag or a robots.txt block does. Orphans just quietly fail to help. On a legacy platform where you can’t predict what a template change touches, “leave it alone and don’t make it worse” is a defensible answer. I’ve seen far more damage done by confident fixes to systems nobody understood than by leaving a known problem in place.

    What I’d do differently now, on any site like that one: stop the bleeding instead of cleaning up the spill.

    Don’t retro-fix eight thousand pages. Change what happens to the next listing when it expires: one rule, one place, testable before it ships. The backlog stays where it is, but it stops growing, and in a year the shape of the site is meaningfully better without anybody having touched the risky part.

    That’s the version of this recommendation I now give whenever the platform is old enough that nobody wants to be the person who broke it.

    The honest caveat

    Removing orphan pages is a cleanup, not a growth lever. I have never seen a site’s traffic jump because orphans were dealt with in isolation, and I would be suspicious of anyone who told you otherwise.

    What it does is stop you wasting effort. On the car marketplace, ninety per cent of the indexed pages were carrying nothing and going nowhere, and every piece of work anybody did on that site, content, links and speed, was being spread across ten times more pages than needed to exist. The value of fixing it is that everything you do afterwards lands somewhere.

    It’s also the fastest way to find out that your real problem is architectural. A site with a 90% orphan rate does not have an internal linking problem. It has a site that was never designed to have a shape.

    The short version

    1. An orphan page is a page nothing links to. It isn’t broken, which is why nothing flags it.
    2. At scale they come from a database, not a writer. Listings, expired products, feeds.
    3. Matching numbers in a crawl report mean one page-set, not several problems. 8,317 orphans and 8,317 bad canonicals were the same pages.
    4. A plain crawl will report zero. You must give the crawler the sitemap as a second URL source.
    5. Ask what you actually care about before you touch anything.
    6. Most of them should be removed, not linked. Some should be linked properly. Some should be consolidated.
    7. It’s a cleanup, not a growth lever, but it’s what makes the growth work land.
    8. On an old platform, fixing the flow beats fixing the backlog. Stop generating new orphans; leave the existing ones if touching them is risky.

    If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling, indexing and architecture layer. Putting the links back where they belong is contextual internal linking. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • What I check first when a site should rank and doesn’t

    A med spa came to me because the site wasn’t showing up locally. Good clinic, real reviews, services people search for every day.

    The address in their footer linked to a Google Maps listing for a different business. So did the address on the contact page. The map embedded on the contact page was that other business too, and the Instagram icon in the footer went to that company’s account.

    Nobody had done anything wrong on purpose. A previous build had been adapted from another clinic’s site and those links came with it. But for anyone arriving, a customer or a crawler, the site’s own contact details pointed somewhere else.

    That took about four minutes to find, and no keyword tool was involved.

    This is what I do when someone sends me a URL and says the site should be ranking and isn’t. I start on the home page and work down. The whole first pass takes up to three days, and almost none of it is spent where people expect.

    Here’s the order, and more usefully, why it’s in that order.

    Before anything: start the crawl

    The first thing isn’t really a check. I put the URL into Sitebulb, and Screaming Frog does the same job, then let it run.

    A crawl on a real site takes time, and there’s no sense watching it. So it churns while I look at the site the way a customer would, and I come back once there’s something to read.

    Worth saying, because published checklists always present themselves as a clean sequence and real work isn’t one. The crawl is the fourth thing I read. It’s the first thing I start.

    1. The home page

    The most important page on the site, and where I always begin.

    I think of the home page as a movie trailer. It isn’t the film. Its job is to show enough of what the business is, snippets, glimpses, a sense of the thing, that a stranger decides to stay. A real photo of the owner. Testimonials from people who exist. Enough of the story that the business reads as legitimate.

    It’s the first impression when you meet someone. You know within seconds whether you’re going to keep talking.

    Does it say what the business actually is? Plainly, near the top, in words a customer would use. A great many home pages describe a mood rather than a business.

    Does it link down to the pages that matter? This is the one I find broken most often. On a local site the home page should link contextually to the service pages, inside the copy, not only from the menu. On a Shopify store, to the collection pages.

    Menu links are fine, but every page has those. A contextual link from the home page body, with anchor text that describes the destination, says something different about what that page is for.

    Two variants of this I’ve found recently. One clinic’s home page linked to the same service page twice, from two different places, and to three other services not at all. Another home page had plenty of text but none of it mentioned a single service by name.

    Is the structure sane? A chiropractic clinic I looked at had seven H1 tags on the home page, several of them empty. There were blank H2 tags too. That’s not somebody’s SEO strategy, it’s a page builder emitting headings for layout reasons.

    I mention this because it’s invisible. The page looked fine. It looked good, in fact. You only see it if you look at the markup, which is why it survives years of people wondering why the site doesn’t rank.

    Is there a testimonial section and an FAQ? Both earn a place.

    I’ll be straight about why this is check one rather than something more technical: a home page that doesn’t establish trust makes everything downstream a waste of effort. Rank a page like that better and you’ve only sent more people somewhere to bounce off. That’s a conversion problem more than a ranking one, but it’s the reason I won’t spend three days on a site and skip it.

    2. The header and footer menus

    Then the navigation, on the assumption nobody has looked at it since the theme was installed.

    I check that links work, where they point, and what’s being given prominence.

    On one clinic site, a header menu item labelled Athlete Care pointed at /elementor-685/. That’s an unfinished page-builder draft. It was in the main navigation of a live site, on every page, being crawled every time.

    On a Shopify store I audited, the header menu item names had all been assigned H2 tags by the theme. So had every product name on the home page and collection pages. A collection page with forty products was handing Google forty H2 headings that were just product names, before any of the actual page content.

    That’s a theme-level default, it’s on thousands of stores, and almost nobody looks for it.

    Elsewhere, a clinic’s blog posts didn’t display the header menu at all. Every post on that site was a dead end. A visitor landing there from search had no way through to a service page except the back button.

    Home, about, contact, blog and the real target pages should be reachable from the menus. Often one of them isn’t.

    3. Reading the crawl

    By now the crawl has finished. This is where the mechanical problems surface: broken pages, multiple or missing H1s, duplicate titles and descriptions, title lengths, canonicals, noindexed URLs, external links, orphan pages.

    Most of that is ordinary and every crawler reports it. Three things are worth more attention than they get.

    Scale. The numbers on a real site are not what people imagine. One Shopify store came back with 1,562 URLs with a missing or empty H1. An auto accessories store had 1,301 URLs missing an H1, and 1,347 pages with either no canonical at all or a canonical still pointing at the http version. The same 1,347 were missing a meta description.

    Nobody creates 1,562 problems by hand. At that scale it’s always one template or one setting, and finding what produced it matters more than the list.

    Pages the platform invented. That Shopify store had 659 tag pages, competing with the collection pages that were meant to rank. Nobody wrote them. Nobody linked to them deliberately. They exist because the platform makes them.

    Orphan pages. These are pages nothing links to. Not from the menu, not from another page, not from anywhere. A visitor can only reach one by typing the URL.

    The auto accessories store had 1,277 orphan pages against 1,630 pages indexed in Google. Most of that site was unreachable from the rest of it.

    That’s the finding I most often have to explain twice, because every one of those pages exists, loads, and looks completely normal when you open it.

    And it isn’t only a big-site problem. Here’s the internal URL summary from a crawl of a small outdoor gear store:

    Sitebulb internal URL report showing 174 total URLs: 164 internal HTML pages, 16 broken, 68 redirects and 60 orphaned

    One hundred and seventy-four URLs in total. Sixty of them orphaned. Sixty-eight redirects and sixteen broken pages on a site you could read end to end in an afternoon.

    Over a third of that store was pages nothing pointed at. The owner knew about none of it, and no amount of writing better product descriptions would have touched it.

    4. Redirects

    Short, and occasionally it answers the whole question.

    The www version should resolve to non-www, or the other way round. I don’t much mind which, as long as it’s consistent. HTTP should redirect to HTTPS. Canonicals should be self-referencing and on the https version, which as above is frequently not the case.

    When this is wrong it’s usually been wrong for years, and it means the site has quietly been running as more than one site.

    Worth knowing how much of this is still out there: HTTP Archive’s 2025 Web Almanac found 91.5% of mobile pages served over HTTPS, so roughly one page in twelve still isn’t, more than a decade after Google called it a ranking signal. But the ones I actually find are almost never a site that never migrated. They are a site that migrated years ago and left the canonicals pointing at http, which no visitor will ever see and every crawler will.

    5. Is there enough site here at all?

    Then I stand back and ask whether there’s enough substance to rank on.

    This is a judgement, and it differs for local, ecommerce and national sites. But I carry rough tripwires:

    • fewer than about 50 well-written blog posts and the site probably hasn’t said enough to be taken seriously on its topic
    • a service or collection page at 300 words with a single heading almost certainly hasn’t covered what a buyer needs. The ones that work tend to run longer, with an H1, several H2s and some H3s underneath
    • target pages generally want an FAQ section

    A medical education site I audited had 189 indexable pages, of which 93 were under 500 words. Half the site was pages that didn’t say enough to be worth indexing.

    One caution on using any of this as a benchmark. The 2024 Web Almanac put the median mobile home page at 364 words. If you measure yourself against the web, 300 words looks perfectly respectable. That is the trap. The median page on the internet is not your competition. The four pages already ranking for your query are, and they are not median, or they would not be there. Clearing the average is how a site ends up indexed, unremarkable, and stuck.

    Now the important part. These are not targets. There is no ideal page length, Google has said so plainly, and writing to a word count produces exactly the padded, say-nothing page that got the site here.

    They’re tripwires. When a page falls well under them, it’s a reliable sign nobody has properly answered the question that page exists to answer. The number tells me where to look. Reading it tells me whether there’s a problem.

    Pad a thin page to 1,200 words and you have a longer thin page.

    6. Internal linking

    I open a handful of pages at random and read the links inside the content.

    What I find, over and over: pages with no contextual internal links at all, pages with far too many, generic anchor text such as “click here” and “read more”, and naked URLs used as anchors.

    Two examples from one med spa. A hormone therapy service page had no internal links whatsoever. A weight loss service page had exactly one, to the contact page. The blog posts linked to each other using the raw URL as the anchor text, and never linked to a service page at all.

    That’s a site publishing content that supports nothing.

    Internal linking is the part of SEO nobody needs permission or budget for. It costs nothing, it’s entirely within the owner’s control, and it’s neglected on most of the sites I’m called into.

    7. The about page

    Almost always thin, and almost always disconnected.

    Two checks: does it link back to the home page, and to contact. Usually neither. One med spa’s about page had a single internal link on it, to the contact page, and nothing else. Another site didn’t have an about page at all.

    The bigger problem is what’s on it. Most about pages say the business was established in a year and is committed to quality. That isn’t a story and it tells nobody anything.

    Every brand has a story. Why the person started, what they did before, who the team are. That’s where the trust the home page promised gets substantiated or doesn’t.

    8. The contact page

    Same pattern. Thin, and treated as a form rather than a page.

    I look for a link back to the home page, the Google Business Profile map embedded, and the address and phone number present as text.

    On local sites this is where the damage tends to be. Beyond the med spa at the top of this post, the recurring one is subtler: the address on the site doesn’t exactly match the address on the Google Business Profile. Not wrong, exactly: a suite number formatted differently, an abbreviation, a missing line. It’s the sort of difference a human wouldn’t notice and a machine can’t ignore.

    Two of the local sites I looked at recently had no GBP map embedded anywhere at all. One had the map embedded, incorrectly, in two places.

    9. Reviews and the profile

    On local work I look at the Google Business Profile alongside the site, because ranking in the map pack and ranking the website are two different jobs and clients rarely separate them.

    Mostly I’m looking at review volume against the competition rather than in isolation. One clinic had 93 five-star reviews, which sounds excellent until you look at who they’re competing with. A med spa had 8. Neither number means anything until you’ve looked at the other businesses in the pack.

    What I don’t look at for the first three days

    You’ll have noticed there’s nothing here about backlinks.

    That’s deliberate. On-page is the foundation of a website, and I want to know the foundation is sound before I look at anything built on top of it. Links pointing at a site whose home page doesn’t say what the business does, whose service pages have no internal links, and a third of whose URLs are orphaned are not going to fix any of that. They’ll send authority into a structure that can’t carry it.

    I do get to off-page work. Comparing a site against the three or four businesses beating it on traffic, ranking keywords, referring domains and pages indexed, is a real part of a full audit and sometimes it’s where the actual answer is. Occasionally a site is doing everything right on-page and is simply outgunned, and the honest finding is that it needs links and time rather than another round of fixes. But I can’t tell that from the outside until I know the site itself is coherent, because otherwise I’m attributing to competition what’s really a broken template.

    So it’s sequencing, not dismissal. Foundation, then everything else.

    The second thing missing is Search Console, and that one’s practical rather than philosophical.

    A proper Search Console audit tells you things nothing else can: what the site is already getting impressions for, which pages Google has actually indexed, which queries sit in positions four to twenty and need a nudge rather than a rewrite. It’s some of the most useful data available on any site.

    But it’s the client’s data, and I need them to grant me access before I can see any of it.

    Which is why the three days above are structured the way they are. Everything on this list can be done from the outside, with nothing but a URL. No logins, no access requests, no waiting on an agency to release something. That’s the entire point of the first pass. It tells me whether there’s a real problem worth either of us spending money on, before anyone has signed anything.

    What happened

    Two of these, a Shopify store and a local clinic.

    Same pattern in both. Nothing moved for three to four weeks after the work went in, then it started. On the local site the service pages began showing up. On the Shopify store the collection pages and the home page both improved in rankings and impressions.

    The three-to-four week gap matters, because that’s the point where most people conclude the work didn’t help. It isn’t long enough to draw a conclusion from. It’s roughly how long it takes for a site to be recrawled, reassessed, and for anything to surface.

    Why I’m not going to claim one change caused that

    Here’s where I part company with how these stories usually get told.

    The convention is a before-and-after graph, a green arrow, and an implied straight line from one change to the recovery. I’m not going to give you that, because I don’t believe it and neither should you.

    Rankings move for many reasons at once. Sites get worked on while other things happen to them: an update lands, a competitor changes, seasonality does what seasonality does, the client runs a campaign nobody mentioned. When a set of fixes goes in together and things improve a month later, the honest description is that things improved after the work, not because of any single item in it.

    Anyone showing you a graph proving one change caused a recovery is guessing with more confidence than the data supports.

    What I can tell you is narrower and more useful: these are the problems that were there, this is the order I found them in, and this is roughly how long before anything moved.

    The short version

    Up to three days, in this order:

    1. Start the crawl and leave it running.
    2. The home page: what the business is, whether it earns trust, whether it links contextually to service or collection pages, and how many H1s it’s actually emitting.
    3. Header and footer menus: working, pointing somewhere real, and not handing out heading tags for layout.
    4. Read the crawl: missing H1s, duplicate and missing titles and descriptions, canonicals, noindexed URLs, platform-generated pages, orphans.
    5. Redirects: one canonical host, HTTPS enforced, self-referencing canonicals.
    6. Is there enough site here: judged by reading, not counting.
    7. Internal linking: contextual links inside the content, anchor text that means something.
    8. About page: linked, substantial, an actual story.
    9. Contact page: linked, map embedded, address matching the profile exactly.
    10. The profile and reviews, on local work, against the competition rather than in isolation.

    All of it from the outside, on nothing but a URL. Backlinks, competitors and Search Console come afterwards.

    Most of the sites I’m called into have a problem somewhere in the first four, and have been spending money on links for a year.


    If a site should be ranking and isn’t, that’s the work I do. An SEO audit is this process run properly and written up. Technical SEO covers the crawling and indexing layer underneath it, including what “crawled, currently not indexed” actually means, which is where a lot of these end up.

  • Crawled, currently not indexed: what it actually means

    Here is the Page Indexing report for a Shopify store I have worked on.

    Search Console page indexing summary showing 479 indexed pages and 6,540 not indexed across 9 reasons

    Four hundred and seventy-nine pages indexed. Six and a half thousand not.

    That is not a broken site. It sells, it ranks, it has a real business behind it. It is a normal Shopify store, and the ratio is normal too, which is the part worth sitting with before you do anything about it.

    Open the same report on your own site and there is a row called Crawled, currently not indexed. On most sites it is not a small number. On a Shopify store it is routinely larger than the number of pages you think you have.

    The status itself is unambiguous. Google found the URL, spent the resource to crawl it, looked at what came back, and decided not to put it in the index. There was no error. Nothing is broken. Google simply did not think the page was worth keeping.

    What happens next is almost always the same. The client sees a number in the thousands, reads it as a verdict on the site, and wants it fixed.

    In my experience that number is mostly pages the platform created on its own. Nobody wrote them. Nobody linked to them. Nobody would miss them. Google has looked at them and declined, which is the correct outcome, and the report is doing its job.

    So before any of the diagnosis below, there is one question that settles most of these conversations: is there anything in that list you actually care about?

    First, find out whether any of it matters

    Do not start at the top of the list and work down. Start by looking for pages that were meant to earn something:

    • the home page
    • service pages
    • collection pages
    • product pages
    • location pages
    • blog posts

    If none of those are in there, you are looking at platform exhaust. The work is a cleanup job, stopping the pages being generated, and not an investigation. A great many clients have spent months worrying about a number that meant nothing was wrong.

    Google’s own crawl budget guidance puts the threshold for even thinking about this at one million or more pages updating weekly, or 10,000-plus pages updating daily. A seven-thousand-URL Shopify store with a few hundred real products is in neither bracket. Whatever is happening on that site, it is not that Google ran out of crawl budget — which is the explanation clients arrive with about half the time.

    If one of those pages is in the list, open it yourself. On a laptop and on a phone.

    Not in a crawler. Not in a tool. In a browser, on both, as a customer would meet it. Mobile matters more than laptop here, because that is predominantly what Google is indexing from. A page that renders fine on a desktop and collapses on a phone is not a mysterious indexing problem. It is a broken page, and you have just found the cause in under a minute.

    A surprising number of these are resolved at exactly that point, because the page is visibly not what the person believed they had published.

    If the page loads properly on both, submit it for indexing in Search Console and watch what happens.

    Then ask whether the page should exist

    Here is the breakdown for that same store.

    Search Console table of reasons pages are not indexed, showing 3,292 excluded by noindex tag, 665 not found, 584 with redirect, 493 blocked by robots.txt, 249 alternate pages with canonical, and 1,241 crawled but currently not indexed

    1,241 pages crawled and not indexed. That alone is more than twice the number of pages Google has actually indexed.

    But look at the row above it. 3,292 excluded by a noindex tag, half the entire report. Somebody, at some point, told Google not to index those. Almost certainly not deliberately, one by one. That is the theme doing it, or an app, or a setting nobody has revisited.

    Most of the URLs in these reports were never deliberately created. They are a by-product of the platform.

    Shopify is the clearest example, because it generates a great deal on your behalf. Collection URLs that duplicate each other. Product URLs available under several paths. Tag pages. Filter combinations. Vendor pages nobody linked to. None of it was a decision anyone made. It is just what the platform does when you add products.

    It is worth knowing how normal that is at scale. Ahrefs looked at 14 billion pages in their Content Explorer index and found 96.55% of them get zero traffic from Google, with another 1.94% getting between one and ten visits a month. Their own caveat is that the sample skews toward the better end of the web. So the honest read is not that 96% of pages are bad — it is that the web is overwhelmingly made of pages nobody was ever going to visit, and a Shopify store generating four thousand filter URLs is not an outlier. It is the median case.

    WordPress does the same thing more quietly. Author archives on a single-author site. Date archives nobody will ever browse. Tag pages with one post in them. Attachment pages, if nothing has been done about them.

    So the first question is not how do I improve this page. It is did I mean to publish this page at all.

    If the answer is no, there is nothing to improve. Google has already made the correct decision about it, and you are looking at a report that is functioning exactly as intended. The work is to stop generating the pages, not to fix them.

    That reframing removes a large share of the report on most sites before you have diagnosed anything.

    When it is a content problem, improve it or delete it

    For the pages you did mean to publish, thin content is the main cause. I do not have a contrarian position on this. The standard advice is correct.

    Where I differ is on the second option. The advice is usually given as improve the content, and improvement is treated as the only acceptable answer. It is not. For a lot of pages the honest answer is that they should not have been written, and the correct action is to remove them.

    A page that exists because someone decided the site needed more pages is not going to become valuable by having more words added to it. Google has looked at it and told you what it thinks. Adding four hundred words does not change the underlying fact that the page has nothing to say.

    Improve the ones with a real reason to exist. Delete the rest. Both are legitimate outcomes, and treating deletion as failure is why these reports grow.

    The causes that are not about content at all

    The remainder are technical, and they are the ones that get missed because everyone has already accepted the content explanation.

    • The sitemap was never submitted. It happens far more often than it should, including on sites that have had an agency for years. It does not directly cause this status, but on a large site with weak internal linking it changes what Google prioritises, and it is a thirty-second check.
    • Googlebot is spending its time on the wrong files. In Search Console, under Settings, the Crawl Stats report breaks requests down by file type. Here is a WordPress site, a local business with fewer than a hundred pages worth indexing:
    Crawl requests by file type: JavaScript 34%, HTML 26%, CSS 26%, other file type 8%, JSON 3%

    JavaScript 34%. CSS 26%. HTML 26%.

    Googlebot spent more of its time on this site fetching stylesheets and scripts than fetching pages. Only about a quarter of everything it requested was a page at all. Google is not refusing to index the content so much as never getting a clear run at it, because most of what it collected was furniture.

    That site is not unusual either. HTTP Archive’s 2025 Web Almanac crawled 16.2 million sites in July 2025 and found the median mobile home page ships 632 KB of JavaScript against 22 KB of HTML — roughly twenty-eight times more script than document. Googlebot has to fetch the first to see the second. What the file-type chart is showing you is not a fault so much as the shape of the modern web arriving in your crawl stats. The question is only whether your site is worse than the median, and by how much.

    That is a plugin and theme problem, not a writing problem, and no amount of rewriting will touch it. It is also two clicks from the report everyone is already staring at, and almost nobody looks at it.

    A warning about reading that chart. If you are looking at a domain property rather than a single host, the breakdown covers every subdomain together. I have a client property showing JavaScript at 95% and HTML at 1%, which looks catastrophic until you open the host list and find that a static asset subdomain accounts for 1.1 million of the 1.16 million requests. The main site is fine. The chart was describing a CDN.

    So segment by host before you conclude anything. The file-type chart is one of the more useful screens in Search Console and one of the easiest to misread.

    The server is not answering reliably. The same screen has a Host Status line covering robots.txt fetching, DNS resolution and server connectivity.

    Crawl stats showing host status reading Host had problems in the past, alongside a crawl breakdown by response code

    Look at what that one says: host had problems in the past. It is a green tick with a caveat attached, and it is very easy to scroll past a green tick.

    Intermittent failures do not always surface as errors anywhere else, and they change how much Google is willing to attempt. Worth opening rather than accepting the tick.

    How long it takes once you have fixed it

    Around three weeks, in most cases, for pages to start being indexed after the actual cause has been addressed.

    Not three weeks from when you make the change. Three weeks from when you make the right change, which is a different date and usually a later one.

    If it has been six weeks and nothing has moved, the assumption to revisit is not the timeline. It is the diagnosis.

    The short version

    The status is not a verdict on your writing. It is a description of what Google did.

    Work through it in this order:

    1. Is anything in the list important? Home page, service, collection, product, location page, blog post. If none of those are there, stop. Nothing is wrong.
    2. If one of them is there, open it on a laptop and a phone. Most problems are visible at this point, and mobile is where they show up.
    3. If it loads properly on both, submit it for indexing and watch what happens.
    4. If the page is one you meant to publish and it is thin, improve it or delete it. Both are fine.
    5. If it is none of those, go technical. Sitemap, crawl distribution by file type, host status.
    6. Give it three weeks from the correct fix, then re-examine the diagnosis rather than waiting longer.

    Most of the sites I am called into have spent months on step four, for pages that never should have got past step one.


    If a site should be ranking and is not, that is the work I do. Technical SEO covers the crawling, rendering and indexing layer. If you are not sure which of several plausible problems is the one costing you, that is what an SEO audit is for.