How do you find duplicate content on a website?

SEO questions ยท Technical SEO

Short answer

Use Search Console’s page indexing report to see which URLs Google treats as duplicates, then a crawler to find pages sharing titles, headings or near-identical content. Most duplication is technical, the same page at several URLs, and is fixed with redirects and consistent canonicals. It is not a penalty. It wastes crawling and splits signals.

Duplicate content causes more anxiety than almost anything else in SEO, mostly because of the word “penalty” attached to it in a thousand old articles. So before how to find it, the reassurance: Google’s own starter guide says duplicate content is inefficient, but it’s not something that will cause a manual action. Copying other people’s content wholesale to manipulate rankings is a different matter. Your own site repeating itself is a housekeeping problem.

The two kinds of duplication

Technical duplication is the same page reachable at more than one URL. http and https, www and non-www, with and without a trailing slash, with tracking parameters, through different category paths, with sort or filter parameters. The content is identical because it is literally the same page.

Content duplication is different pages saying nearly the same thing. Forty location pages with the town name swapped. Product descriptions copied from the manufacturer. Tag and category archives listing the same posts. Two blog posts written years apart on the same topic.

They need different fixes, so it helps to find them separately.

Finding technical duplicates

  1. Search Console page indexing report. Look for “Duplicate without user-selected canonical”, “Duplicate, Google chose different canonical than user” and “Alternate page with proper canonical tag”. These are Google telling you directly which URLs it considers copies.
  2. URL inspection. For a page you care about, check which URL Google selected as canonical. If it is not the one you intended, something is sending mixed signals.
  3. A crawler. Crawl the site and look for identical titles or content hashes across different URLs. Most desktop crawlers flag exact duplicates automatically.
  4. Test the variants manually. Type your homepage with and without www, with http, and with a trailing slash. Each should redirect to one version.

Finding content duplicates

Crawlers can flag near-duplicates by similarity score, which is useful for templated pages like locations or products. Sort by similarity and read the top pairs. For older blog content, export your posts with their main target query and look for two posts chasing the same thing.

Search Console’s performance report helps too. Filter by a query and look at the pages tab. If two of your pages both receive impressions for the same query and trade places over time, they are competing. That is cannibalisation, a close cousin of duplication.

For content copied from elsewhere, search Google for an exact sentence from your page in quotes. If other sites carry the same text, you are not the original source in Google’s eyes, or you are sharing it with many others.

Fixing what you find

Google lists the main consolidation methods by strength. Redirects are a strong signal, rel=canonical is a strong signal, and sitemap inclusion is a weak one. In practice:

For technical variants that should never exist, such as http or non-www, redirect them sitewide.

For variants that need to exist for users, such as filtered or sorted views, use a canonical pointing at the main version. Then make sure internal links and the sitemap point at the same version. A canonical contradicted by your own links is a weak signal.

For content duplicates, decide which page should win. Improve that one, then merge the others into it and redirect, or make each genuinely different. Forty thin location pages are usually better as six real ones.

Mistakes I see

Canonicalising every page in a paginated series to page one. It sounds tidy and it is against Google’s guidance, which says to give each page its own canonical. I admit in four ways canonical tags get misused that I recommended this myself in older audits.

Using robots.txt to hide duplicates. It stops Google crawling them, which means it cannot see your canonical or noindex, and the URLs can still appear in results.

Panicking about boilerplate. A shared header, footer and sidebar on every page is not duplicate content in any sense that matters.

How much duplication is a problem

A few variants on a small site barely matter; Google usually picks the right version. It becomes a real problem at scale: stores where filters generate thousands of near-identical pages, publishers with tag archives for every keyword, sites where one template error creates a duplicate of every page. That is where crawling time is wasted and signals are split badly enough to cost rankings.

A twenty-minute check you can do today

Open the page indexing report and note the counts for the three duplicate statuses. Inspect your homepage and your three most important pages to confirm Google chose the canonical you intended. Then type your domain four ways, with and without www and https, and confirm each one lands on the same address. If all of that is clean, duplication is probably not what is holding the site back, and your time is better spent elsewhere.

If Search Console is full of duplicate statuses and you are not sure which ones matter, my technical SEO service finds the source, usually a template or a parameter, and fixes it once.

More questions about technical SEO

See all 5 questions about technical SEO or browse every SEO question I answer.