Skip to content
Have a project in mind?
The Gloria JournalSEO

Duplicate content: what it really costs your rankings

Wooden rubber stamp reading Duplikat beside a blue ink pad, with its imprint on white paper

The same text reachable at two addresses, a product page lifted from the supplier's catalogue, an article copied by a competitor: three situations, one label — duplicate content — and one recurring question. Will Google penalise me for this? Google's own documentation answers in a sentence: some duplicate content on a site is normal, and it does not breach its spam policies.

That does not make duplication harmless. It costs visibility, it wastes crawling time, and it sometimes surfaces the wrong URL. What follows sets out the real risk, case by case, and the three situations worth your attention: content copied by a third party, supplier product descriptions, and machine translation.

Canonical tags, redirects and URL parameters are a separate job, usually an output of a technical SEO audit rather than a decision taken page by page.

What Google actually does with duplicate content

Google does not remove duplicates from its index. It groups them. Pages whose main content is identical or very close are clustered, and Google designates the one it judges most complete and most useful. That version appears in the results; the others are crawled less often.

You can state a preference — a redirect, a rel="canonical" tag, inclusion in the sitemap — but Google is explicit that these are signals, not instructions, and that it may select a different URL (Google Search Central, page updated 20 August 2026).

The consequences are therefore economic, not punitive. Google gives four reasons to handle the matter yourself: choosing which URL is shown, consolidating accumulated signals — incoming links included — on one address, keeping your analytics readable, and not wasting crawling time. On a brochure site the stake is low. On a catalogue of several thousand URLs it becomes measurable.

The duplicate content SEO penalty: what actually exists

No Google document describes a penalty for duplicate content. What the spam policies do sanction are two specific behaviours, both of which assume intent and volume.

  • Scraped content: taking content from other sites, often automatically, and hosting it to manipulate rankings. Google cites republishing without original value, copying with synonyms substituted, and reusing feeds with no benefit to the reader.
  • Scaled content abuse: generating many pages for the primary purpose of manipulating rankings. The examples include scraping feeds or other content to produce pages at volume, "including through automated transformations such as synonymising, translation or other obfuscation techniques".

Both rules sit in the Google spam policies, updated on 28 August 2026. The Manual Actions report in Search Console carries no heading called duplicate content either: duplicates can only fall under a broader category, thin content with little or no added value.

One order of magnitude on the algorithmic side. After the March 2024 core update, Google announced on 26 April 2024 that it had reduced low-quality, unoriginal content in its results by 45% (Google blog, 2024). That is an internal estimate with no published methodology: a direction of travel, not an independent measurement.

Working out whether this affects you

Three checks are enough to size the problem, and they stop you rebuilding a site to solve something marginal.

  • The Page indexing report in Search Console. Two statuses relate to duplicates: "Duplicate without user-selected canonical" and "Duplicate, Google chose different canonical than user". Their volume tells you whether the subject is marginal or structural.
  • A crawler such as Screaming Frog, to find identical titles, meta descriptions and H1s. These are not offences in themselves; they are symptoms.
  • A quoted search for a full sentence from one of your pages, to see who else has published it.

A recurring pattern: a site is relaunched, every page ends up with the same title and meta description, and catalogue URLs read like index.php?cPath=1_8. Identical tags are not a violation. They are the sign of a template published without customisation, and neither Google nor a human can tell those pages apart in a results page. Fix what is visible first — titles, descriptions, H1s — and leave the URLs until later, because changing addresses always costs something, even with a clean 301 redirect. Our note on how to plan an SEO audit covers that sequencing.

When someone else copies your content

This is the case that worries people most and does the least damage. Scraped content is a violation on the copier's side, not yours. Google groups the two versions and generally keeps the original, which it has known about for longer and which links point to. It does not always work: a highly visible site that republishes a page from a rarely crawled site can get ahead of it. Three sensible reflexes.

  • Establish the facts before acting. Run the quoted search, then use URL Inspection in Search Console to see which canonical Google selected. If your page is still the version shown, there is nothing to do.
  • Reinforce precedence. Request indexing as soon as you publish, link the page from your other content, and keep it in the sitemap. Where new pages are slow to be picked up at all, the real issue is usually indexing rather than duplication.
  • Ask for removal. If the copy infringes your copyright, Google provides a copyright removal request form. The procedure removes the URL from its results; it does not delete the page from the site hosting it, which is a separate legal matter.

Deliberate syndication is a contractual question rather than a technical one. If you allow a partner to republish your articles, the handling happens on their side — a canonical pointing to your page, or noindex — and that clause is negotiated when the agreement is signed, not afterwards.

E-commerce: supplier product descriptions

Retailers routinely resell the same references as dozens of competitors, using the same supplier text. Reusing the manufacturer's description is not sanctioned. The difficulty is elsewhere: Google has no reason to prefer your page to two hundred others showing the same paragraph. It will keep one of them, and not necessarily yours.

On a catalogue of several thousand references, rewriting everything is unrealistic. The trade-off is made on revenue and search volume: the references that sell get their own content, the rest keep the supplier text. What genuinely differentiates a product page is your own photographs, the details the manufacturer left out, real-world uses, and the questions your customer service team keeps answering.

Variants, facets and pagination

Shop-specific cases generate most internal duplication. Size and colour variants produce near-identical pages; filter combinations multiply URLs nobody searches for; paginated listings repeat the same introduction throughout. None of it attracts a penalty, but all of it spends crawling time on pages you would never choose to rank. Decide which version you want indexed, express that with a canonical or a redirect, and keep the rest out of the sitemap. Where the answer is to strengthen a page rather than hide it, that is a content problem rather than a technical one.

Machine translation: the most exposed case

This is the only form of duplication explicitly named in Google's spam policies, as one of the automated transformations that add no value. Translating an existing text mechanically creates nothing for the reader.

The phenomenon is documented. In A Shocking Amount of the Web is Machine Translated (Thompson et al., ACL Findings 2024), the authors analyse 6.38 billion sentences drawn from web crawl data: 57.1% belong to groups of translations available in at least three languages, and measured quality falls as that number rises. Note the scope: sentences already identified as translations, not the web as a whole.

In practice, a machine translation reviewed and corrected by a human raises no problem: it is a production tool, not a spam strategy. A site duplicated into twelve languages at the click of a button falls into the category Google says it wants to reduce — and carries the separate commercial risk of inaccurate product or legal wording. If you cannot review the output, publish fewer languages; our page on multilingual WordPress sites sets out what that involves.

Summary: which risk for which case

Duplicate content: Google's handling and the action to take, case by case
Situation How Google handles it Real risk Action
Internal duplicates (filters, parameters, http/https) Grouped, Google picks the URL Diluted signals, wasted crawling Redirect or canonical
Supplier description reused as is No sanction, page is undistinctive Weak visibility against other retailers Enrich the references that matter
Content copied by a third party Violation on the copier's side Rare: the copy outranks you Check, then removal request if copyright applies
Authorised syndication Grouped across domains Your page is not the one shown Canonical or noindex on the partner's side
Unreviewed machine translation Covered by the spam policies Lost visibility, possible manual action Review, or reduce the number of languages
Mass copying of third-party content Covered (scraped content) Manual action possible Do not do it

The useful question is therefore not "do I have duplicate content?" — every site produces some — but "who decides which page is shown, and what do my pages offer that cannot be found elsewhere?".

Common questions

Is there a Google penalty for duplicate content?

No. Google's documentation states that some duplicate content on a site is normal and does not breach its spam policies. What is sanctioned is scraped content and the mass production of pages with no added value.

At what similarity percentage does content count as duplicate?

Google publishes no threshold. It groups pages whose main content is identical or very close, then designates one as canonical. The practical test is whether two pages deserve to appear separately in a results page.

What should I do if another site copies my articles?

Check in Search Console whether your page is still the version Google shows; in most cases the original is kept. If the copy infringes your copyright, Google provides a removal request form that takes the URL out of its results, without deleting the page itself.

Can I use my supplier's product descriptions?

Yes, this is not a violation. But a description identical to every other retailer's gives Google no reason to choose your page. Prioritise rewriting on the references that generate revenue or search demand.

Does Google prohibit machine translation?

Not as a tool. The spam policies target the generation of many pages through automated transformation, translation included, with no added value for the reader. A translation reviewed and corrected by a human raises no difficulty.

Two pages on my site are very similar: should I delete one?

It depends on their use. If the duplicate no longer serves visitors, a permanent redirect to the page you keep is the cleanest answer. If it still needs to be reachable, a canonical tag tells Google which version is the reference.

This link opens in a new tab.