Author: Mehul Dedhia

  • Most of your product pages cannot rank

    Most product pages cannot rank, and telling a client that is more useful than another list of on-page recommendations.

    Not because the pages are badly built. Because there are three thousand of them, they are competing against Amazon and the manufacturer, and the effort available is enough to properly do about twenty. The question worth answering is which twenty.

    The base rate is worth having in front of you before that conversation. Ahrefs looked at 14 billion pages and found 96.55% get zero traffic from Google, with a further 1.94% getting between one and ten visits a month. Product catalogues are a large part of what makes that number what it is.

    The filter I use

    Four tests, in this order. A page has to clear all four.

    • 1. Does anybody search for this specific thing? Not the category. The product. Some products have named demand: a model number, a brand-plus-type phrase, a thing people know they want. Most have none, and no amount of copy creates it. If the only query that could reach this page is the category term, the category page should be ranking, not the product.
    • 2. Can you say something the manufacturer’s page does not? If your description is the manufacturer’s description, and it will be on most catalogues, you are asking Google to prefer your copy of a page over the original. It will not. Either you have fitting notes, comparisons, real photographs, sizing advice from actually handling the thing, or you have nothing to offer.
    • 3. Is the product going to exist in six months? Ranking takes months. Investing in a page for a line that turns over seasonally is spending the effort in the wrong place. That effort belongs on the category, which persists.
    • 4. Would you back it with internal links? If you would not link to it from the category intro and the home page, you do not really believe it can rank. That is a useful thing to notice before writing eight hundred words.

    What happens to the other 2,980

    They still need to exist, be indexable, and convert when someone lands on them. They just are not ranking projects.

    What they need is the cheap, template-level baseline: a unique title, a real H1, correct structured data, an image with alt text, and a breadcrumb linking up to a category that is a ranking project. That is one template change covering all of them, which is a very different job from writing three thousand descriptions.

    And a portion of them should not exist at all. Discontinued lines, variants that are really the same product, one-off items from four years ago. Thin at scale is usually not a writing problem. It is a decision nobody made about which pages deserve to be pages.

    The mistake this replaces

    The standard recommendation is “add 500 words to every product page.” I have written a version of that myself, in older audits, and I would write it differently now.

    It is not wrong so much as unactionable, and unactionable advice does real damage: the client hears three thousand descriptions, concludes it is impossible, and does nothing. Meanwhile the twenty pages that could genuinely have ranked stay untouched, and the template baseline that would have helped all three thousand never gets done either.

    Twenty good pages plus one template fix beats three thousand mediocre descriptions, and it is a job somebody can actually start on Monday.

    Where the category page fits

    On most catalogues the category is the page that can rank and the product is the page that converts. That split is worth being explicit about, because it changes where the writing goes.

    A category page with real copy, contextual links to its best products, and a genuine answer to the question the category term implies will outperform the sum of its thin product pages. It is also one page rather than four hundred.

    And the competition on the product end is structural rather than something you out-write. HTTP Archive’s 2025 Web Almanac found ecommerce software on 19.2% of mobile sites, so a very large number of stores are listing versions of the same manufacturer-supplied text about the same items. Being the four-hundredth copy of a description is not a position improved by adding two hundred words to it.

    The short version

    1. Most product pages cannot rank, and saying so is more useful than another optimisation list.
    2. Four tests: named demand, something the manufacturer does not say, longevity, and whether you would link to it.
    3. Everything else gets the template baseline: title, H1, schema, alt text, breadcrumb.
    4. Some of them should not exist.
    5. “Add 500 words to every product page” is unactionable, so nothing happens at all.
    6. Categories rank, products convert. Put the writing where the ranking is.

    If a site should be ranking and it isn’t, that’s the work I do. Internal linking is how you back the pages you have chosen. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • Click depth is a symptom, not a rule

    Click depth is how many clicks it takes to get from the home page to a given page. The received wisdom is that everything should be within three, and like most received wisdom it is roughly right and badly explained.

    Three is not a rule. What is actually true is narrower: depth is a symptom, and the number tells you how much of your site the internal linking has abandoned.

    Why depth matters at all

    Not because Google counts clicks and applies a penalty at four. Because internal links are how importance moves around a site, and a page five clicks down is receiving whatever is left after four divisions.

    The practical consequence is that deep pages get crawled less often and rank worse than their content deserves, and on a large site they may not get crawled at all. The extreme version of this is a page nothing links to, at depth infinity, which is the case where 8,317 pages out of 9,220 were unreachable.

    So depth and orphaning are the same problem measured differently.

    The number that actually matters

    Not the maximum depth. The distribution.

    A site where 90% of pages sit at depth two or three and a handful sit at six is fine. The six are probably archives. A site where the bulk of the catalogue sits at depth five has an architecture problem, and the fix is structural rather than a matter of adding links.

    Any crawler will give you this as a chart. It takes ten seconds to read and it is one of the more useful ten seconds in an audit.

    For a sense of what dense linking looks like, the 2024 Web Almanac put the median page across the whole web at 41 internal links, rising to 129 on the top thousand sites by traffic. Depth is the other side of that number: a site with few links per page has to be deep, because there is nowhere else for the pages to hang.

    What actually causes depth

    • Pagination. The commonest cause by a distance. If a category has twenty pages of products and each page links only to the next, product 400 is twenty clicks from the category. Everything on page eight onwards is functionally invisible. This is why “load more” and deep pagination are an architecture decision, not a UX one.
    • Too few categories. Counterintuitive, but a flat structure with three categories and two thousand products creates depth, because the only way to reach most products is through pagination. More categories, properly chosen, makes the site shallower.
    • Navigation that hides the middle. A menu that links to top-level categories and a footer that links to policies. Everything in between is reachable only by browsing.
    • Content that never got linked. Blog posts from four years ago, reachable only through the archive. Not harmful, but they are not doing anything either.

    On a genuinely large site depth stops being theoretical, because crawling becomes rationed. Google’s crawl budget guidance puts that threshold at a million pages updating weekly or ten thousand changing daily. Below it, deep pages are crawled less often but they are crawled. Above it, the deep end of the site can simply stop being visited.

    What I do about it

    Fix the categories before adding links. If depth comes from pagination, the answer is a better category structure so there is less to paginate. Adding a “popular products” block to the home page is a plaster on a structural problem.

    Make the hubs real. Category pages with copy and contextual links pull their children up a level. This is the same fix as most internal linking problems.

    Do not solve it with a sitewide links block. Adding a footer with two hundred links technically reduces depth to two and achieves nothing, because a link that appears on every page says nothing about the relationship between any two pages.

    Accept depth for things that deserve it. Old blog archives, legal pages, individual reviews. Not everything needs to be three clicks from the home page, and pretending otherwise produces the two-hundred-link footer.

    The honest caveat

    I have never seen a site’s traffic move because click depth was reduced in isolation. It is a diagnostic more than a lever. A deep site is telling you something about how the architecture was designed, and it is that underlying thing you are fixing.

    If someone has handed you “reduce click depth” as a headline recommendation with no explanation of what is causing it, ask what is causing it.

    The short version

    1. Three clicks is not a rule. Depth is a symptom of internal linking, not a metric to hit.
    2. Read the distribution, not the maximum.
    3. Pagination is the commonest cause. Product 400 on page 20 is invisible.
    4. Too few categories creates depth, not the other way round.
    5. Fix the structure, not the symptom. A 200-link footer games the number and changes nothing.
    6. Depth and orphan pages are the same problem measured two ways.

    If a site should be ranking and it isn’t, that’s the work I do. Internal linking is the layer underneath this one. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • Internal linking when the catalogue is too big to do by hand

    The advice in most internal linking posts is: open the page, find a relevant phrase, link it to the page you want to rank. That works, and I recommend it, and it does not survive contact with a catalogue of nine thousand products.

    Nobody is hand-placing links on nine thousand pages. So the question on a large site is different: what structures produce good internal links automatically, and which of the automatic ones you already have are working against you?

    Start by auditing what the templates already do

    On a large site, almost every internal link on the page was placed by a template, not a person. So the first job is not adding links. It is reading what the templates are already saying.

    Take one product page and count where its links go. Typically: the main navigation, a breadcrumb, a “related products” block, a “recently viewed” block, the footer. That is the entire vote that page casts, repeated across the whole catalogue.

    Now ask what those blocks are ranked by. Related products is usually “same category, random” or “same category, newest”. Nobody chose that. It means every product on the site is passing its links to an arbitrary handful of siblings, and the pages you actually want ranking are getting no more than the ones you do not.

    The four structures that do the work

    Breadcrumbs. The highest-value automatic internal link on an ecommerce site and the cheapest to get right. Every product links up to its category, every category to its parent. That is a clean hierarchy expressed in links, generated for free, and a surprising number of sites either lack them or have them pointing at a flat structure that does not reflect the categories.

    Category and collection pages, treated as real pages. These are the hubs. If a category page is a grid with no copy, it has nothing to link out from except the grid. Two paragraphs of intro copy on a category page is not a content exercise. It is a place to put contextual links to the subcategories and the two or three products that matter most.

    Related items, ranked deliberately. Change “random within category” to something with intent behind it: best sellers, higher margin, the products you have decided to compete on. This is one template change that reweights the entire catalogue.

    The blog pointing at the catalogue. On most stores the blog links to the blog and nothing else. A closed circuit. Every post should link to at least one page that sells something, in the body, with descriptive anchor text.

    What the aggregate data suggests

    In the 2024 Web Almanac, the median page across the whole web carried 41 internal links. On the top thousand sites by traffic the median was 129.

    I would not read that as cause and effect, since large sites have more pages to link to. But the direction matches what I find: the sites that are winning have dense, deliberate internal linking, and the sites I get called into have whatever the theme shipped with.

    What I would not do

    Automated keyword-to-link plugins. They produce links nobody chose, with anchor text that repeats site-wide, and they will eventually link a word in a customer review to a category page. The value of an internal link is partly that somebody decided to put it there.

    The automated checks will not catch the resulting mess either. The 2025 Web Almanac found 86.6% of mobile home pages pass Lighthouse’s descriptive-link-text audit, but that test catches “click here” and nothing else. It has no opinion on a site where every link is generated, repeated site-wide, and pointing at whatever the plugin matched.

    • A sitewide block of forty links. If every page links to forty things, none of them is being emphasised. This applies to fat footers especially.
    • Ignoring the orphans. Before any of the above, run the crawl with the sitemap fed in and find out how many pages have nothing pointing at them. On one catalogue that number was 8,317 out of 9,220. No amount of template tuning matters while nine pages in ten are unreachable.

    The short version

    1. At scale, links are placed by templates, not people. Audit the templates first.
    2. Read one product page’s outgoing links, because that is the vote every page is casting.
    3. Breadcrumbs are the cheapest high-value structure.
    4. Category pages need copy so they have somewhere to link from.
    5. Rank “related items” deliberately instead of randomly. One change, whole catalogue.
    6. Point the blog at things that sell.
    7. Find the orphans before any of this. Nothing else matters if the pages are unreachable.

    If a site should be ranking and it isn’t, that’s the work I do. Internal linking is the cheapest link building there is. This is the version for sites too big to do it by hand. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • Faceted navigation: index, crawl or block, per facet

    Faceted navigation is the single largest generator of URLs on most ecommerce sites, and the decision about it is almost never made. It is inherited from the theme and then discovered two years later in a crawl.

    The maths is the problem. Six filters with five options each is not eleven pages. It is every combination anyone can click, and the number runs into the tens of thousands on a catalogue that has a few hundred products in it.

    The wrong question

    “Should we index faceted URLs?” has no answer, because it treats every facet as the same thing. They are not.

    Some facets describe how people actually search. “Waterproof walking boots” is a query with volume behind it, and if you sell waterproof walking boots then that filtered view is a legitimate landing page, arguably a better one than the category above it.

    Other facets describe how the warehouse is organised. “Sort by newest”, “12 per page”, “colour: seven shades of grey”. Nobody searches those and nobody ever will.

    So the decision is per facet, and it takes an afternoon with somebody who knows the products. That is the actual work and there is no shortcut through it.

    It is worth doing because of how much of the web this describes. HTTP Archive’s 2025 Web Almanac found ecommerce software on 19.2% of mobile sites, with Shopify at 25.3% of those and WooCommerce at 44.4%. Faceted filtering ships as standard on all of them, which means the URL explosion is the default state and the tidy version is the exception somebody had to choose.

    The three buckets

    • Index it. The facet matches real demand, the resulting page has enough products on it to be useful, and you are willing to write an intro and a title for it. Usually this is one or two facets: brand, or a core product attribute. Treat those views as category pages, because that is what they are.
    • Crawlable but not indexed. Let Google follow the links so it can reach the products, but keep the filtered view out of the index with a noindex. This is the right default for most combinations, and it is the bucket people forget exists. The choice is not binary between “index” and “block”.
    • Not crawlable at all. Sort orders, pagination sizes, session parameters, anything combinatorial. These should not be links Google can follow.

    Why noindex rather than robots.txt

    This is the mistake I see most and it is worth being precise about.

    Blocking a URL in robots.txt stops Google fetching it. It does not remove it from the index, and if the URL is already indexed, blocking it means Google can never read the noindex you have also added, so the page stays in there indefinitely with no snippet.

    Worse on a faceted site: blocking the filter paths can cut off the route to products that are only reachable through a filter. You have then made products invisible in order to tidy up a report.

    HTTP Archive’s 2025 Web Almanac found meta robots tags on 47.9% of mobile pages but a noindex directive on only 2.4%. The mechanism that should be doing this work is barely used.

    The tell that you have a problem

    Compare three numbers: products in the catalogue, pages you think you have, and pages Google has indexed. On the sites where these diverge badly, facets are usually most of the gap.

    Facets are also a common source of pages nothing links to, because a filtered view can be generated, indexed, and then dropped from the navigation when the filter set changes.

    The other tell is in the crawl: filter parameters appearing in URLs that also carry no canonical, or a canonical pointing at themselves rather than the unfiltered category.

    What I would do first

    Not a facet strategy document. Count them. Export the URL list, group by parameter, and see which facets are producing the volume. On most sites two or three parameters account for nearly all of it, and dealing with just those gets you most of the benefit for a fraction of the argument.

    Then decide those two or three properly, and set the rest to the safe default.

    The short version

    1. Facets are combinatorial. Six filters is not six pages.
    2. The decision is per facet, not per site. Some match real demand; most describe the warehouse.
    3. Three buckets: index, crawlable-but-noindexed, not crawlable.
    4. The middle bucket is the default and the one people forget exists.
    5. Use noindex, not robots.txt. Blocking makes indexed pages permanent and can hide products.
    6. Count before you strategise. Two or three parameters usually cause nearly all of it.

    If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling, indexing and architecture layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • Log files tell you what Googlebot actually did

    Every crawler tells you what is on your site. Log files tell you what Googlebot actually did about it, and those are not the same document.

    This is the one technical check I rate highly and recommend rarely, because for most sites the effort is not worth it. I want to be straight about that before explaining what it is good for.

    What only logs can tell you

    • Which pages Googlebot has never fetched. Not “not indexed”. Never requested at all. A crawler tells you a page exists and is indexable. A log tells you Google has not been near it in three months. Those pages are invisible in every other report because there is no data about them to show.
    • How often it comes back. Crawl frequency by section is a rough proxy for how much Google cares about each part of the site. If the blog is fetched daily and the service pages monthly, that tells you where the perceived value sits, and it is often the reverse of where the business wants it.
    • What it is wasting requests on. Parameter URLs, old redirects, files nobody links to any more. You can infer this from the Crawl Stats report in Search Console, but the logs give you the actual URLs rather than a category chart.
    • Whether it is really Googlebot. A meaningful share of traffic claiming to be Googlebot is not. Verifying by reverse DNS changes the numbers, sometimes a lot.

    When I actually ask for them

    Three situations, and outside these I do not raise it.

    A large site where pages are not getting indexed and nothing else explains it. If a section is sitting in Discovered, currently not indexed at volume, which is a different status from pages nothing links to at all, logs answer whether Google is ignoring it or has never found a path to it. That is a different fix each way.

    After a migration. Logs show whether Google has picked up the new URLs and how fast it is abandoning the old ones. Nothing else shows the transition in progress.

    When somebody claims a crawl budget problem. Logs settle it in an afternoon, usually by showing there is no such problem. which there usually is not.

    Google’s own threshold for the question being worth asking starts at a million pages updating weekly, or ten thousand changing daily. If the site is not near either, logs will confirm what the thresholds already told you, which is a legitimate reason to pull them when a client insists, and not a reason to start there.

    Why I recommend it rarely

    Getting the logs is the hard part and it has nothing to do with SEO. On managed hosting they may not be retained. Behind a CDN you need them from the CDN as well or you are looking at a fraction of the traffic. Someone has to be persuaded to hand over a file that also contains visitor IPs, which is a data protection conversation before it is a technical one.

    Then the volume is awkward, enough to be annoying but not enough to justify tooling, and the analysis needs a person who knows what a normal pattern looks like.

    For a local business or a store with a few hundred products, all of that effort buys you very little. The Crawl Stats report in Search Console covers the same ground well enough, for free, in two clicks. I would exhaust that first every time.

    What I look at when I do have them

    Requests by response code. A high share of 404s or 301s means Google is spending its time on URLs that no longer exist. That is the most common finding and the easiest to act on.

    Requests by directory. Where the attention is going, at a glance, and whether it matches where the money is.

    Requests by file type. Scripts and stylesheets against actual pages. I have seen sites where the majority of Googlebot requests were not pages at all.

    That last one is less surprising than it sounds once you look at what a page now weighs. In HTTP Archive’s 2025 Web Almanac, drawn from 16.2 million sites, the median mobile home page ships 632 KB of JavaScript against 22 KB of HTML. Googlebot has to fetch the furniture to see the document, so a crawl log dominated by assets is partly just the modern web. The question is whether your ratio is worse than that, and by how much.

    Pages in the sitemap with zero requests. The most useful single query in the whole exercise, and the one you cannot run any other way.

    The short version

    1. Logs tell you what Googlebot did. Crawlers tell you what exists.
    2. Only logs show pages Google has never fetched.
    3. Worth it for large sites with unexplained indexing gaps, migrations in progress, and disputed crawl budget claims.
    4. Not worth it for most small and mid-size sites. Crawl Stats covers it.
    5. Getting the file is harder than reading it, and it is a privacy conversation first.
    6. Sitemap URLs with zero Googlebot requests is the query that justifies the exercise.

    If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling and indexing layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • 9,220 pages indexed. Nobody decided to publish them.

    A site with 9,220 pages indexed and roughly 8,000 visits a month. Another with 1,630 indexed and a catalogue of a few hundred real products. A third with 1,265 indexable pages, of which 1,082 were blog posts.

    None of those businesses set out to publish that much. In every case the platform did it, one URL at a time, over years, and nobody was counting.

    That is crawl bloat: the gap between the pages you decided to publish and the pages that exist. It is not a penalty, it is rarely a crawl budget problem, and it is one of the most common things I find.

    Where the extra URLs come from

    The sources are boringly consistent across platforms.

    Expired or discontinued items that keep their page. The car marketplace above was a listings site. Vehicles came and went, which is the business, and the pages stayed, unlinked, in the sitemap, indefinitely.

    Filter and facet combinations. Every combination is a URL. The maths is combinatorial, so a store with six filters does not have six extra pages, it has hundreds.

    Archives nobody browses. Author archives on a single-author site. Date archives. Tag pages with one post on them. On the fitness site, author archives ran several pages deep for authors with a handful of posts each.

    Pagination without canonicals. 266 paginated URLs on that same site, each one a page in Google’s eyes.

    URL segments that add nothing. 1,082 posts carrying /blog/ for no reason is not bloat in the count sense, but it is the same inattention showing up in a different column.

    Why it matters, stated honestly

    Not because Google punishes large sites. It does not.

    It matters for a duller reason: everything you do afterwards gets divided by it. Internal links, content effort, whatever authority the domain has: all of it spreads across the pages that exist rather than the pages that matter. On the marketplace, ninety per cent of indexed pages were unreachable, carrying nothing and going nowhere, and every piece of work anyone did on that site was being diluted ten to one.

    The second reason is diagnostic. A site where nobody has counted the URLs is a site where nobody has been looking at the technical layer at all, and the rest of the audit usually confirms it.

    It also helps to know the base rate before you panic about a number. Ahrefs examined 14 billion pages and found 96.55% get zero traffic from Google. Pages existing and earning nothing is the normal condition of the web, not a symptom peculiar to your site. What makes bloat worth acting on is not that the pages are idle. It is that you are dividing everything you do by them.

    What it is not, on most sites, is a crawl budget problem. Google’s own threshold for worrying about that starts at a million pages, or ten thousand changing daily. A 9,220-page site is not close. The budget explanation is available, comforting, and usually wrong.

    Counting yours

    Three numbers, ten minutes.

    What Google has indexed. Search Console, Page indexing, the indexed count.

    What you think you have. Ask the person who runs the site. Do not look it up. The discrepancy between belief and reality is itself the finding.

    What a crawl finds, with the sitemap fed in as a second source so orphans surface.

    If those three numbers are within sight of each other, you do not have this problem. If the first is several times the second, you do.

    What I actually recommend

    • Stop the flow before clearing the backlog. This is the recommendation I have changed my mind about. Retro-fixing eight thousand URLs on an old platform is risky and slow. Changing what happens to the next expired listing is one rule in one place, testable before it ships. The backlog stops growing, and in a year the shape of the site is meaningfully better without anyone touching the part nobody understands.
    • Use noindex, not a robots block. Blocking a URL means Google cannot read the noindex on it, so the page stays indexed forever. This is backwards from what most people reach for first.
    • Take them out of the sitemap. A sitemap listing pages you do not want indexed is you actively submitting them.
    • Decide per facet, not per site. Some filtered views are real landing pages people search for. Most are not. That is a judgement per facet and it is the only part of this that takes real thought.

    The short version

    1. Crawl bloat is the gap between pages you published and pages that exist.
    2. Expired items, facets, archives and pagination produce almost all of it.
    3. It is not a penalty. It dilutes everything you do afterwards.
    4. It is usually not a crawl budget problem, because Google’s threshold starts at a million pages.
    5. Count three numbers: indexed, believed, crawled. The gaps are the finding.
    6. Stop generating before you clean up. Safer, cheaper, and it holds.
    7. Noindex, not robots.txt. Blocking makes pages permanent.

    If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling, indexing and architecture layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • An SEO audit is not a crawl report

    If what you were handed is a list of issues sorted by severity, with a count next to each, you were handed a crawl report. It may have cost you two thousand pounds and it may have a cover page with your logo on it. It is still a crawl report.

    I am not being sniffy about tools. I use them on every job and the good ones are excellent. The distinction I care about is narrower and it decides whether the document is worth anything: a crawl tells you what is on the site. An audit tells you which of those things is the reason you are not ranking.

    Those are different claims and only one of them requires somebody to take a risk.

    Three things a crawl cannot do

    It cannot tell you whether a page should exist. A crawler will report a thin page. It has no opinion on whether the page was a mistake, and “improve this page” and “delete this page” are both correct answers to the same row depending on a judgement no tool makes.

    It cannot connect two findings. On one site I audited, the crawl reported 8,317 orphan pages and, separately, 8,317 pages with a missing or http canonical. Two rows, two counts, one page-set, one template. Read as a list it is two projects. Read as a diagnosis it is one cause and one fix.

    It cannot tell you what the business needs. A tool will flag a missing H1 on a legal disclaimer and a missing H1 on the main service page with the same severity. Only one of those matters and the tool has no way to know which.

    The volume problem is real and measurable. WebAIM’s automated crawl of the top million home pages averaged 56.1 detected errors per page. Meanwhile the 2025 Web Almanac found 86.6% of mobile home pages pass Lighthouse’s descriptive-link-text audit, so the same tooling that hands you fifty-six problems per page will also tell you your link text is fine when every link on the site comes from the menu. Loud where it should be quiet, silent where it should not.

    What I think the difference actually is

    An audit commits to a cause.

    That is the whole of it. A crawl report lists everything that is true. An audit says: of all these true things, this is the one costing you, and here is why I think so. That sentence can be wrong, which is exactly why most deliverables avoid writing it. A list of four hundred issues cannot be wrong. It also cannot be acted on.

    The test I would apply to any document you have paid for: does it name a cause, and would the author look foolish if that cause turned out to be something else? If nobody has taken that risk, nobody has done the diagnosis.

    What should be in it that a tool cannot produce

    • Somebody opened the site. On a laptop and on a phone. A surprising share of problems are visible in the first four seconds and invisible to a crawler, which does not care that the home page describes a mood rather than a business.
    • Search Console, read properly. Not exported. Read. Which queries have impressions and no clicks, what the average position column actually says, whether the pages you care about are in the index at all.
    • A comparison against the two or three businesses beating you. Sometimes the honest finding is that the site is fine and simply outgunned, and that finding is worth more than forty rows of fixes.
    • A prioritised three. Not sixty. Three, named, with who does each one.
    • Things you should not bother with. An audit that recommends everything has not made any decisions. The items I explicitly rule out are usually the most useful page in the document, because they are the ones saving you money.

    Why so many deliverables are crawl reports

    Because they are safe and they look substantial. Four hundred rows feels like value. Three findings and a paragraph explaining what you should ignore feels thin, right up until you try to implement either one.

    There is also a straightforward incentive problem: a list of four hundred items can be produced in an afternoon and defended forever. A diagnosis takes a day of reading and can be proved wrong in six weeks.

    The short version

    1. A crawl says what is on the site. An audit says what is costing you.
    2. A crawl cannot judge whether a page should exist, and improve/delete are both valid answers to the same row.
    3. It cannot connect findings. Matching counts mean one cause; the tool cannot see that.
    4. The test: does the document name a cause the author could be wrong about?
    5. It should include somebody actually opening the site, on a phone.
    6. It should tell you what not to do. That page is usually the most valuable one.

    If a site should be ranking and it isn’t, that’s the work I do. What I check first is the sequence I run before any tool is involved. If you want somebody to name the cause rather than list the symptoms, that’s what an SEO audit is for.

  • How to read a crawl without drowning in it

    The first time you export a crawl of a real site you get a spreadsheet with forty columns and several thousand rows, and the honest reaction is to close it.

    I have been doing this for fifteen years and I still read almost none of it. Here is the order I actually go in, which takes about twenty minutes and answers most of the question before the detailed reading starts.

    Before you crawl: give it the sitemap

    This is the step that changes the output more than any setting in the tool.

    A crawler starting from the home page follows links. By definition it will never reach a page nothing links to. So on a site with eight thousand orphan pages, a plain crawl reports zero orphans and you get a clean bill of health that is completely wrong.

    Feed it the XML sitemap as a second URL source and let it compare the two lists. Anything in the sitemap the crawl never reached is an orphan. On one marketplace that comparison produced 8,317 pages nothing on the site linked to, out of about 9,220 indexed. None of them would have appeared otherwise.

    The five things I look at, in order

    • 1. Total indexable URLs versus what you think you have. One number against the client’s expectation. If they say “about two hundred pages” and the crawl says 1,630, stop and find out what is generating the difference. That gap is often the entire finding and you can have it in ninety seconds.
    • 2. Status codes, grouped. Not each one. How many 200s, 301s, 404s, 5xx. What I want is the shape: a site with 68 redirects out of 174 URLs has a redirect problem; a site with three has housekeeping. Any 5xx at all gets looked at immediately.
    • 3. Indexability, and why not. Every crawler has a column for this. Sort by the reason. Noindex, canonicalised elsewhere, blocked. What you are looking for is a page you care about sitting in a bucket you did not intend.

    Two figures worth carrying while you read that column, both from the 2025 Web Almanac: a canonical tag is present on just 67% of mobile pages, and 10.3% have an invalid element inside the <head>. The second one matters more than its position on any issues list, because the head closes at the first invalid element, so a canonical or a meta robots sitting below the break is not being read at all, while your crawler cheerfully reports it as present.

    4. Word count, sorted ascending. Crude and I use it every time. It surfaces the empty templates, the placeholder pages, the category pages nobody finished. On one platform this showed 466 of 812 pages under the threshold, and the useful part was not the count, it was that they were all the same page type.

    Resist benchmarking that against the web. The 2024 Web Almanac put the median mobile home page at 364 words, so almost anything you find will look respectable by comparison. The median page on the internet is not your competition. The four already ranking for your query are, and they are not median or they would not be there.

    5. Crawl depth. How many clicks from the home page. Anything important sitting at depth five needs an explanation, and usually the explanation is that it is only reachable through pagination.

    That is it for the first pass. Titles, headings, meta descriptions and alt text: all of that is real and none of it is where I start, because none of it explains why a site that should rank does not.

    The habit that saves the most time

    Look for numbers that match.

    When two findings report the same count, they are describing the same pages. 8,317 orphans and 8,317 bad canonicals are not two problems. 1,347 pages missing a canonical and 1,347 missing a meta description are one page-set with one template behind them.

    I check for this deliberately now, before reading anything in detail, because it collapses a long report into a short list of causes. A four-hundred-row report is usually six to twelve actual problems.

    What the crawl will not tell you

    It will not tell you whether a page should exist, whether the home page says what the business is, or which of the fifty flagged issues is the one costing money. Those need somebody to open the site and form an opinion.

    Which is why I run the crawl in the background and start by looking at the site myself. The crawl is corroboration. It is not the diagnosis.

    The short version

    1. Feed the crawler your sitemap or it cannot find orphans, and will report zero.
    2. Compare total indexable URLs to what the client believes. The gap is often the whole finding.
    3. Read status codes as a shape, not row by row.
    4. Sort indexability by reason and look for pages you care about in the wrong bucket.
    5. Sort word count ascending. Crude, fast, reliable.
    6. Matching counts mean one cause. Check for them before reading anything closely.
    7. The crawl corroborates a diagnosis. It does not make one.

    If a site should be ranking and it isn’t, that’s the work I do. What I check first is the sequence the crawl runs alongside. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.

  • If you only do three things, which three?

    Every audit I hand over has more in it than the client will ever do. That is not a flaw in the audit. It is what an audit is, a complete picture, and the complete picture has always been bigger than one quarter’s capacity.

    So the useful question is not what is wrong. It is: if you only do three things, which three?

    Here is how I answer it, and why the answer is rarely the top of the list.

    What the frequency data says

    I went back through eleven audits written since 2021 and counted what actually came up.

    • Internal linking: 11 of 11. Every single one.
    • Schema markup missing or incomplete: 10 of 11.
    • Thin content: 10 of 11.
    • Heading tags used decoratively: 8 of 11.

    After that it falls away quickly. That distribution is the starting point for prioritising, but it is not the answer, because frequency is not severity. Internal linking is in every audit partly because it is genuinely always broken and partly because it is easy to spot.

    The web-wide picture says the same thing about how little these counts discriminate. HTTP Archive’s 2025 Web Almanac found an H1 present on only 70% of mobile pages, and WebAIM’s crawl of the top million home pages found an average of 56.1 detected accessibility errors per page. If you ranked by prevalence you would spend the year on things almost every site has and almost no site is losing because of.

    The three questions I rank by

    1. Is anything actively removing pages from search? A stray noindex, a robots block on a live section, a canonical pointing at the wrong place, a server that returns 5xx under crawl. These go first regardless of how few pages they touch, because they are the only category that is subtracting. Everything else on the list is a failure to add.

    Most audits have none of these. When one is present it is the whole first slot.

    2. What can be fixed once and apply everywhere? A template fault beats a page fault every time. One theme file that adds a missing H1 to 1,562 URLs is a better use of a slot than rewriting nine service pages, even though the nine pages are more visible and feel more like progress.

    This is also the test that stops a list of four hundred rows being a list of four hundred jobs. Group by cause and it usually collapses to six or twelve.

    3. Which fix does the client actually control? This is the one that gets left out of prioritisation frameworks and it decides more outcomes than the other two.

    A recommendation that needs a developer who is booked until November is not a priority. It is a plan. If the person in front of me can edit pages but not templates, then the three things have to be things they can do, or nothing happens and the audit joins the others in a drawer. I would rather ship the second-best fix in a fortnight than the best one never.

    What I put last, and say so

    • Schema, despite being in ten of eleven audits. It is real, it is worth doing, and it almost never changes anything on its own. It sits near the bottom unless the site is competing for a rich result that still exists.
    • Meta descriptions at scale. Twenty pages, yes. Thirteen hundred, no.
    • Heading level tidiness. Fixing menus wrapped in H2 is worth doing at template level. Reordering H3s under H2s across a blog is not.

    How I write it now

    I stopped delivering a ranked list of everything. It reads as a menu and people pick the easy items.

    Instead: three things, named, with who does each one and roughly how long. Then everything else in a second section clearly marked as later. The client can argue with three. Nobody can argue with sixty, so they do not. They just do nothing, which looks like agreement and is not.

    The short version

    1. Frequency is not severity. Internal linking was in 11 of 11 audits; it is rarely the most damaging thing.
    2. Anything actively removing pages from search goes first. It is the only category that subtracts.
    3. Template fixes beat page fixes. One change, every URL.
    4. Rank by what the client can actually do, not by what is theoretically most valuable.
    5. Schema, bulk meta descriptions and heading order go last, and I say so out loud.
    6. Deliver three, not sixty. A list of sixty is a list of none.

    If a site should be ranking and it isn’t, that’s the work I do. What I check first is the sequence behind these priorities. If you’re holding a report you can’t prioritise, that’s what an SEO audit is for.

  • Deindexed pages: find out what shipped that week

    A page that was in Google and is now not is a different problem from a page that was never accepted, and the two get treated as one thing constantly.

    If the page was never indexed, you are asking whether it is good enough, which is the crawled, currently not indexed conversation and it is a question about quality. If it was indexed and dropped out, something changed. Those are different investigations and the second one is usually faster, because a change has a date attached to it.

    Start with what changed, not with the page

    This is the part people skip. They open the page, read it, decide it is fine, and then start rewriting it anyway.

    Before touching the content, get the date it dropped from Search Console and ask what happened that week. A release. A theme update. A plugin. A migration. A robots.txt edit. A new SEO plugin with opinionated defaults. On the sites I get called into, the answer is in that list far more often than it is in the writing.

    I have seen a whole section vanish because a staging Disallow shipped with a release, and the client had spent three weeks improving the content on those pages first.

    The order I check

    Is it accidentally noindexed? View the raw source, not the rendered page, and search for noindex. Then check the HTTP headers, because X-Robots-Tag does the same job invisibly and almost nobody looks there. An SEO plugin with a category-level setting can do this to a thousand URLs without anyone touching a page.

    The numbers say how rarely this is deliberate. The 2025 Web Almanac found meta robots tags on 47.9% of mobile pages but an actual noindex directive on only 2.4%. Half the web emits the tag; almost nobody uses it to make a decision. So when a noindex does turn up on a page that was ranking, the overwhelming likelihood is that a setting produced it rather than a person.

    Is it blocked in robots.txt? If so, and it also carries a noindex, the two are fighting. Blocking wins and removal loses, because Google cannot read a directive on a URL it is not allowed to fetch.

    Does the canonical point somewhere else? A page canonicalising to another page is asking to be dropped. This is the one that surprises people, because the page looks completely normal.

    Did the URL change? A slug edit with no redirect deindexes the old URL and starts the new one from nothing. On a CMS this is a two-click accident.

    Is it still linked from anywhere? A page that quietly became an orphan during a redesign has stopped being part of the site, and Google eventually treats it that way.

    Did the server have a bad week? Sustained 5xx errors during a crawl will do it. Crawl Stats, Host status.

    When none of that is it

    Then, and only then, it is a quality judgement, and it usually arrives with company: a group of pages rather than one, often after a core update.

    I would be honest with yourself at that point about whether the page was ever earning anything. Ahrefs looked at 14 billion pages and found 96.55% get no traffic from Google at all. A page dropping out of the index after years of no clicks is not a mystery to solve. It is the index catching up.

    The pages worth fighting for are the ones that were doing something. For those, the work is the same work as any thin page: make it the best answer to something, or accept it was not going to be.

    What I would not do

    • Mass re-submission. Requesting indexing on four hundred URLs does not fix a cause and it does not speed anything up. Submit one, watch what happens, and use it as a test rather than a treatment.
    • Rewriting first. If the cause is a noindex header, a month of content work changes nothing and you will conclude that content does not matter.

    The short version

    1. Dropped out is not the same as never accepted. Get the date first.
    2. Ask what shipped that week before you read the page.
    3. Check noindex in the source and in the headers. X-Robots-Tag is invisible on the page.
    4. Then robots.txt, canonical, URL changes, internal links, server errors, in that order.
    5. Only then treat it as quality, and be honest about whether the page was ever earning.
    6. Do not mass-resubmit. It treats a symptom and tells you nothing.

    If a site should be ranking and it isn’t, that’s the work I do. Technical SEO covers the crawling and indexing layer. If you’re not sure which of several plausible problems is costing you, that’s what an SEO audit is for.