SEO and Organic Search

What Is Duplicate Content on a Website? (And How We Actually Dealt With It)

Diagram showing three URLs serving the same content consolidating into one canonical page that Google shows in search results

Most explanations of duplicate content are written in the abstract. Here’s a concrete one: while rebuilding this website, we found 59 old blog posts on the previous contentkings.ie, most following the same “10 Tools To X” / “5 Ways To Y” template, covering largely the same handful of content-marketing topics with barely any variation between them. That’s duplicate content in practice, not theory - and dealing with it was part of this site’s own migration, not a hypothetical example.

What duplicate content actually is

Duplicate content is substantially identical or near-identical content that exists on more than one URL - either across different websites, or (the more common problem) across different pages of the same website. It’s not usually about someone copying your article wholesale. More often it’s:

  • The same product or service described on multiple near-identical pages (one per location, one per slight variant)
  • A blog platform serving the same post at /post-title/, ?p=123, and a tag/category archive page that repeats the full content
  • www and non-www, or http and https, versions of the same page both being indexed
  • Printer-friendly or “amp” versions of a page that duplicate the main content
  • A content library that’s grown for years without anyone checking whether old articles still say anything new

Does it actually hurt your SEO?

Google has been fairly consistent on this: duplicate content isn’t typically a “penalty” in the sense of punishment for wrongdoing. What actually happens is more mundane and, in a way, more wasteful - when multiple URLs carry the same content, Google has to choose one to show in search results and effectively ignores the rest. Any links or authority pointing at the “losing” URLs get diluted rather than consolidated. Your site also burns crawl budget on pages that add nothing, which matters more the larger a site gets.

The exception is content scraped or duplicated with genuinely deceptive intent (keyword-stuffed doorway pages, content farms), which Google’s spam policies treat more seriously. Most real-world duplicate content on legitimate small business sites isn’t that - it’s accidental, structural, or just old content nobody’s revisited.

How we actually found and fixed it

For our own migration, the process was:

  1. Inventory everything - we pulled both the live WordPress database and the full Wayback Machine archive (1,029 historical URLs going back to 2011), because relying on just the current database missed years of content that had already been deleted from the live site.
  2. Classify, don’t assume - each piece of content got assessed on its own merits: genuinely useful and distinct, templated filler with no first-hand experience behind it, or outright theme boilerplate.
  3. Decide per-URL, not in bulk - the 59 templated posts got a 410 (Gone), not a blanket redirect to the homepage. A few genuinely good pages (the About page, the real service descriptions) got rewritten and kept. Nothing was redirected into the new site just because it used to exist.

That’s the practical version of fixing duplicate/thin content at the source, rather than trying to patch it after the fact with canonical tags alone.

What to actually do about it

  • Canonical tags (<link rel="canonical">) tell search engines which version of a near-duplicate page is the “real” one - essential for www/non-www, pagination, and URL parameter variants, but not a fix for content that’s genuinely thin or repetitive.
  • 301 redirects consolidate genuinely duplicate URLs into one.
  • A proper content audit, not just a technical crawl, is the only way to catch the content-mill problem specifically - pages that are technically unique URLs but say nothing meaningfully different from each other. A crawler won’t flag that for you; reading the content will.
  • Don’t be afraid to retire content. Keeping a thin, duplicate-adjacent page live because deleting things feels wasteful is usually the wrong call. A smaller site made entirely of pages worth reading beats a larger one padded with near-duplicates.

If you’re looking at an older website and suspect this might be a problem, that’s exactly the kind of thing a proper content audit or technical SEO review is for.

← Back to Insights