Duplicate content issues and how to fix them (without breaking your rankings)
A client sent me a message last month that made me wince: "Google Search Console says 1,847 pages are 'Duplicate, Google chose different canonical than user.' What do I do?"
That number wasn't a disaster. It was a symptom. And once I dug into the crawl, the real problem was embarrassingly simple — a mix of trailing slashes, UTM parameters bleeding into the index, and a plugin that had quietly generated three copies of every product page.
Duplicate content is one of those topics where everyone has an opinion and almost nobody has a decision tree. You'll read "use a canonical tag" twenty times without learning when a canonical is the wrong fix. That's what this article is about.
Key Takeaways
- Duplicate content rarely triggers a penalty. The real damage is diluted link equity and wasted crawl budget.
- Canonical tags are a hint, not a directive. If you need certainty, use a 301 redirect.
- Fix the source, not the symptom. Rewriting content is the last resort, not the first move.
- On faceted e-commerce sites, parameter handling in Search Console often solves 80% of the problem.
- Give Google 3–6 weeks after fixes before judging results. Reindexing doesn't happen overnight.
- Similarity thresholds matter: I flag pages above 85% textual overlap, not identical copies.
What does "duplicate content" mean?
Strictly speaking, duplicate content is any block of content that appears at more than one URL — either as an exact copy or as a near-match. The problem isn't the duplication itself. It's that search engines don't know which version to trust, so they pick one, and sometimes they pick wrong.
Two flavors exist, and people constantly confuse them.
- Internal duplication — you create it yourself through URL structure, CMS quirks, or faceted navigation.
- External duplication — someone else copies your content, or you syndicate it to partners and marketplaces.
- Near-duplicates, the sneaky middle ground: same page with two paragraphs swapped, or a product description with only the color changed.
The 25% reality check
Years ago, Google's Matt Cutts publicly stated that roughly a quarter of all content on the web is duplicate — and that most of it isn't spam. That framing still holds up. Google's own documentation repeats the same idea: duplicate content isn't a penalty trigger in itself, it's a ranking and consolidation problem.
Which means when someone tells you "I got penalized for duplicate content," they almost certainly didn't. They got outranked by a better-consolidated version of the same page.
Can Google penalize you for duplicate content?
No — and this is where most blog posts get it wrong by hedging. Google does not apply a duplicate content penalty in the classic sense. There's no filter that drops your site because two pages look alike.
What actually happens is less dramatic and more expensive.
Google splits your signals. If page A and page B both target "trail running shoes," and half your backlinks point to A while the other half point to B, neither page accumulates enough authority to rank. I've watched a client sit at position 14 for a keyword for eight months because of exactly this — a www/non-www split nobody had noticed. Consolidating the two versions took that keyword to position 4 in about five weeks.
The second cost is crawl budget. Every duplicate URL Googlebot fetches is a fetch it didn't spend on a page that matters. Sites under 500 URLs rarely feel this. Sites over 50,000 feel it painfully.
The one case where you can get hit
Deliberate manipulation. Scraping someone else's content and republishing it, spinning articles, or generating doorway pages purely to capture search traffic. That's not duplication as a technical accident — that's spam, and it's treated as such.
How to diagnose the problem before you fix anything
Running tools before understanding your duplication pattern is how people end up with 40 redirect chains and a broken sitemap. Diagnose first.
Step 1: check what Google already knows
In Search Console, open the Page indexing report. Filter for Duplicate, Google chose different canonical than user and Duplicate without user-selected canonical. Export both lists. That export is your actual to-do list — not a generic audit recommendation.
Step 2: crawl your own site
Screaming Frog is still my default here. Point it at your domain and look for three things: near-duplicate titles, near-duplicate content (the near-duplicate detection feature), and URL patterns that shouldn't exist. I also run Siteliner on smaller sites when I want a fast similarity percentage without configuring a full crawl.
My working threshold: pages above 85% textual similarity get flagged. Below that, I usually leave them alone. Chasing 60% overlap is a rabbit hole with no payoff.
Step 3: find inbound duplicates
Copy a distinctive sentence from your best-performing page and search it in quotes. If a scraper site ranks above you for your own text, that's a different problem — and it's usually solved by filing a DMCA request rather than touching your own markup.
How do you fix duplicate content?
There is no single fix. There's a decision, and the decision depends on whether the duplicate URL is one you want to keep.
The decision tree I actually use
| Situation | Fix | Why |
|---|---|---|
| Two live URLs, one is clearly better | 301 redirect | Certain, permanent, passes link equity |
| Both URLs need to stay live (params, print versions) | Canonical tag | Non-destructive; Google treats it as a hint |
| Low-value page that shouldn't rank at all | noindex | Removes it from the index without a redirect |
| Faceted navigation with hundreds of variants | Parameter handling + canonical | Scales; manual rules would take weeks |
| Content genuinely exists on a partner site | Cross-domain canonical | Tells Google which version is the original |
| Near-duplicate but both have traffic | Rewrite or merge | Only option when neither should disappear |
Notice what's missing from that table: "canonical everything." Canonicals get ignored constantly — when they point to a noindexed page, when they conflict with hreflang, when they're injected by JavaScript that renders too late. If you need certainty, redirect.
The http/https and www traps
These are the cheapest wins in the entire discipline. Pick one protocol, one hostname, and force everything else to it with a server-level 301. Then make sure your internal links, sitemap, and canonical tags all use that same version.
I've seen a site lose roughly a third of its organic entries because every internal link still pointed to the http version three years after migrating to https. Nobody noticed because the redirect worked. But every single internal link was a redirect hop — slow for users, wasteful for crawlers.
Handling syndicated content
If you publish on a partner site or a marketplace, agree on one thing before anything else: who gets the canonical. Usually it's you, with a link back and a cross-domain canonical from them. Some partners will refuse. In that case, delay the syndicated publication by 7–14 days so your version gets indexed first.
I'll admit I underestimated this for years. A B2B client lost a head term to a partner who republished their guide the same day, and it took four months of outreach to get a canonical added.
International duplication and AI content pitfalls
Two edge cases trip up almost everyone, and both have gotten worse recently.
hreflang is not a canonical
If you have an English and a French version of the same page, that's not duplicate content — that's localization, and hreflang is the correct signal. The mistake I keep seeing is combining hreflang with a canonical pointing to the English version. That combination effectively tells Google the French page shouldn't exist. Don't do it. Each language version should be self-canonical.
AI-generated content and duplication
Machine-written pages don't create duplication in the classic sense, but they create something similar: dozens of near-identical articles that differ only in phrasing. Google's guidance is consistent — helpful, original content is what matters, regardless of how it was produced. The practical effect is that generic AI output tends to compete with your own pages for the same queries.
When I audited a content site last year that had published around 200 AI-assisted articles, the pattern was obvious: 31 pages were cannibalizing each other on the same cluster of keywords. Merging them into six well-researched pieces roughly doubled the cluster's total sessions in about two months. Fewer pages, more traffic. That's the usual outcome.
How long do fixes take to work?
Longer than you want, and the delay is uneven.
- 301 redirects: consolidation usually shows within 2–4 weeks on mid-sized sites.
- Canonical tags: often 4–8 weeks, and Google may take months to switch its chosen canonical.
- noindex: removal from the index typically happens within days, but signals transfer is slower.
- Content merges: expect 6–12 weeks before positions stabilize.
Request reindexing in Search Console for the URLs you changed. Don't submit the same URL repeatedly — it doesn't speed anything up and it clutters your history.
The thing most people forget
After you redirect a URL, update every internal link pointing to it. Don't leave redirect chains behind. A redirect that resolves in one hop is fine. A redirect that chains through three hops is a crawl-budget leak you'll never find until you run a crawl report six months later and wonder why your crawl depth looks broken.
Duplicate content isn't exciting work. It's plumbing. But plumbing is what keeps a site's rankings from quietly leaking away — and the fix is almost never "write more content." It's deciding, page by page, which version deserves to exist. Most sites have about a dozen of those decisions to make. Once you've made them, the rest of your SEO work stops fighting itself.