On a small ecommerce site with fifty products, duplicate content is a minor housekeeping matter. On a large one, with tens of thousands of SKUs, filters, variants and parameters, it becomes one of the most damaging and least visible problems you can have. A store with 10,000 products and a handful of filter options can expose hundreds of thousands of URLs to search engines, most of them near-identical, and the risk is you can fragment your ranking signals, waste crawl budget on pages that will never earn traffic, and leave search engines guessing which version of a page deserves to rank.
The reassuring news is that duplicate content is almost always fixable, and fixing it can produce some of the largest gains available in technical SEO. This is the checklist we work through when we tackle it on large e-commerce stores, and the order we tackle it.
First, clear up the biggest myth
Before the checklist, one thing worth settling, because it causes a lot of unnecessary panic. There is no such thing as a duplicate content penalty. Google does not hand out a punishment for having duplicates.
What actually happens is more subtle, and still costly. When several URLs carry the same or very similar content, search engines have to pick one to treat as the authoritative version, and your ranking signals, links, relevance and authority, get split across all the copies instead of concentrated on one. That fragmentation is what harms rankings, along with the crawl budget wasted on redundant URLs that could have been spent on your important pages. So the goal is not to avoid a penalty. It is to consolidate signals onto one clear version of each page.
The checklist
1. Audit the scale of the problem first
You cannot fix what you have not measured, and on a large site the true number of indexable URLs is almost always far higher than owners expect. Start by crawling the site with a tool like Screaming Frog or Sitebulb to find duplicate and near-duplicate pages, then cross-reference against Google Search Console. In the Search Console coverage report, the statuses to look for are “Duplicate without user-selected canonical” and “Alternate page with proper canonical tag,” which tell you how Google is currently handling your duplicates. Compare the number of URLs Google has discovered against the number of products you actually sell. If Google has found ten times more URLs than you have products, you have found your problem.
2. Fix URL hygiene at the source
A surprising share of duplicate content comes from small inconsistencies that generate multiple addresses for the same page. Mixed upper and lower case in URLs, inconsistent trailing slashes, and http versus https or www versus non-www variants all create technical duplicates. Enforce a single lowercase, consistent URL format at the server level, and make sure every variant redirects to your one canonical format. This is unglamorous work, but it removes an entire category of duplication before you touch anything more complex.
3. Apply self-referencing canonicals to every product page
Every product page should carry a canonical tag pointing to its own clean URL. This matters most where the same product is reachable through several paths, such as tracking parameters, affiliate codes or a version nested under a collection. The self-referencing canonical tells Google to treat the clean URL as the master copy and to ignore the parameter-laden variants, consolidating all their signals onto the one page you want to rank. On many platforms this is handled by default, but on a large site it is always worth verifying rather than assuming.
4. Get faceted navigation under control
This is the single biggest generator of duplicate content on large stores, and it is where most of the work usually lies. Filters for size, colour, price and so on create a new URL for almost every combination, and because a shopper can select the same filters in a different order, you often get multiple different URLs serving identical results. Left unmanaged, this alone can spawn hundreds of thousands of near-duplicate pages.
The strategy is selective indexing. Decide which filtered views deserve to be indexed and which do not, based on real search demand. A filtered page like “black walking boots” may warrant its own indexable page because people search for it and it returns a distinct set of products. An obscure combination of four filters that nobody searches for does not. For the pages you want to keep out of the index, apply NOINDEX tags or block the parameters from being crawled, and point canonicals from low-value filter combinations back to the main category page. This is nuanced, situational work with no single rule that fits every case, which is exactly why it repays doing carefully. We treat this as closely related to the wider job of planning a website structure that search engines can navigate cleanly.
5. Handle product variants deliberately
Products that come in multiple colours or sizes are a common source of near-duplicate pages, since each variant page is largely identical to the others. Decide on an approach rather than letting the platform sprawl: often the cleanest option is a single product page that handles variants on the page itself, or a canonical from each variant to a primary version. What you want to avoid is a dozen almost-identical pages competing with each other for the same search.
6. Replace duplicated manufacturer descriptions
Not all duplicate content is internal. When you use the manufacturer’s product description, so do all the other retailers selling that item, which makes your page one of many identical copies across the web. Search engines then fall back on domain strength to decide who ranks, which rarely favours smaller stores. Rewriting descriptions for your priority products with original, useful copy is the fix. It is a content job rather than a technical one, but it belongs on the same checklist because the effect is the same: differentiation instead of duplication. This is worth building into your work on collection and product page content.
7. Deal with pagination and empty pages
Two loose ends round out the list. For paginated category pages, make sure your approach lets Google crawl through to deeper products without treating every page as a duplicate of the first. And for placeholder or empty category pages, those with no products or almost no content, apply a NOINDEX tag until they are populated with useful content, so they do not sit in the index as thin duplicates dragging on the site’s overall quality.
8. Reinforce it all with consistent internal linking
Finally, make your internal links agree with your canonicals. Every internal link should point to the primary, clean version of a page, never to a parameter-laden or filtered variant. Inconsistent internal linking undermines even a well-planned canonical strategy, because it sends search engines mixed signals about which version you actually consider authoritative. Getting this right reinforces everything above.
The checklist at a glance
Work through these in order, since the early steps remove problems that would otherwise complicate the later ones.
| Check | What to do | Why |
|---|---|---|
| ☐ | Crawl the site with Screaming Frog or Sitebulb and cross-reference Google Search Console coverage. Compare indexable URLs against your actual product count. | Reveals the true scale of duplication before you start. If Google has found far more URLs than you have products, that gap is your problem. |
| ☐ | Enforce one lowercase, consistent URL format at server level and redirect all variants (case, trailing slash, http/https, www) to it. | Removes an entire category of technical duplicates created by small inconsistencies. |
| ☐ | Add a self-referencing canonical to every product page pointing to its clean URL. | Consolidates signals from parameter, affiliate and collection-nested variants onto the one page you want to rank. |
| ☐ | Apply selective indexing to faceted navigation: index filter views with real search demand, noindex or block the rest, canonical low-value combinations to the category page. | Faceted navigation is the biggest duplicate generator on large stores. Selective indexing stops hundreds of thousands of near-duplicate URLs forming. |
| ☐ | Handle product variants with a single page or a canonical from each variant to a primary version. | Prevents a dozen near-identical colour or size pages competing for the same search. |
| ☐ | Rewrite manufacturer descriptions with original copy on priority products. | Stops your pages being identical to every other retailer’s, so you rank on differentiation rather than domain strength alone. |
| ☐ | Set pagination so Google can crawl deeper products, and noindex empty or placeholder category pages until populated. | Keeps thin and duplicate pages out of the index where they drag on site quality. |
| ☐ | Make every internal link point to the primary, clean version of each page. | Reinforces your canonicals instead of sending search engines mixed signals about the authoritative version. |
Why this is worth the effort
Duplicate content sits in an unusual position. It is invisible to most site owners, it carries no dramatic penalty, and yet resolving it can lift organic traffic substantially, because you are taking signals that were scattered across thousands of redundant URLs and concentrating them where they belong. There is a further reason it matters more than ever: AI search depends on retrieving accurate, consistent content, and when duplicate or outdated versions of a page exist, they weaken your whole content footprint and make it less likely your preferred page is the one retrieved and cited. Clean, canonical, consolidated pages help you across both traditional and AI-driven search at once.
Because these issues creep back in as a catalogue grows, this is not a one-off fix but something to audit regularly, ideally quarterly and after any significant site change. If you would like us to audit your store for duplicate content, or build the fixes into a wider SEO services and technical SEO strategy, get in touch with the team at TAL and we will help you concentrate your signals where they count.

We’d love to chat
The best ideas start with a good old conversation. Let’s have a chat about how we can help you.
Complete the form and one of the directors will be in touch.
Or just pick up the phone and call us, we’re on 03334049888.













