Blog
-
How to read a crawl without drowning in it
Export a crawl of a real site and you get forty columns and several thousand rows. After fifteen years I still read almost none of it. Twenty minutes in the right order answers most of the question. Feed the crawler your sitemap first or it cannot find orphans and will report zero.
-
If you only do three things, which three?
Every audit I hand over has more in it than the client will ever do. That is what an audit is. So the useful question is which three things. Internal linking was in eleven of eleven audits and is rarely the most damaging. Frequency is not severity. Deliver three, not sixty.
-
Deindexed pages: find out what shipped that week
A page that dropped out of Google is a different investigation from a page that was never accepted, and the two get treated as one thing constantly. Get the date first, then ask what shipped that week, and ask the developer rather than the marketer. Quality comes last, not first.
-
Soft 404s: Google is usually right
A soft 404 is a page that returned 200 and then failed to contain anything. People treat it as a bug in Google’s classifier. In my experience it is almost always correct. Ask whether the URL should exist before asking how to clear the flag, and do not pad the page to satisfy a tool.
-
Crawl budget: almost nobody reading this has one
Almost nobody reading this has a crawl budget problem. Google publishes the thresholds: a million pages changing weekly, or ten thousand changing daily. Crawled, currently not indexed is not a budget problem, it is Google reading the page and declining. You cannot raise it. You can stop spending it.
-
The 1,630-page store with no robots.txt, and what that actually told me
A car parts retailer with 1,630 pages indexed, selling every day, returned a 404 on robots.txt. That is not a crisis. Thirteen per cent of sites do it. What it told me was that nobody had ever been into the technical layer, so I went and looked at what else had been left.
-
1,562 missing H1s and seven H1s on one page. Same issue, opposite problem.
One store had 1,562 URLs with a missing H1. One page had seven. Any crawler files those under the same heading. An issue count measures how far a template reaches, not how much it costs you. Four hundred rows is usually six to twelve real problems once you divide by cause.
-
Four ways canonical tags get misused, and one of them is my own advice
Four sites, four things done wrong with one tag, and the fourth is a recommendation I made twice and got wrong. I told two clients to canonicalise paginated pages to page one. Google says give each page its own. Missing canonicals at scale are one template mistake, not eight thousand.
-
The client never knew I existed, and that was the point
The best result I can show from the last two years belongs to a law firm I have never spoken to. An agency brought me in white label, I set the strategy and did the link building, their team implemented. The agency is my client and decides what the end client hears.
-
Two things I recommended in most of my audits, and have stopped recommending
I counted what I actually recommend across eleven audits since 2021. FAQ schema appeared in seven and nofollowing outbound links in eight. I would make neither now. One is out of date because Google changed something. The other was never right, and I repeated it for years because everyone repeats it.










