Near-duplicate pages
duplicate_contentShort answer
These pages are nearly identical to another page on the site (a SimHash distance of 6 or less) and had no Google impressions in the last 28 days. There is no duplicate-content penalty, but Google chooses one version and the others carry no weight. If a page has no separate purpose, merge it (301) or point its canonical to the main version; if it should rank for something of its own, make its content distinct. Pages that do appear in search are not listed.
Why it matters
Google's guidance is that duplicates are consolidated, not punished — but consolidation means only the chosen URL ranks, and links or content invested in the others are wasted. Typical sources: print versions, tag/category pages with the same post lists, parameter variants, location pages with the city name swapped, product variants as separate URLs. On large sites these pages also eat crawl budget.
How Glimana detects it
Glimana computes a SimHash of each page's main text and compares pages pairwise; pairs within Hamming distance 6 are "similar" (3 or less is a copy). Only indexable pages with zero impressions in the last 28 days are reported, so a short page that ranks on its own is never flagged. The evidence for each page lists its near-duplicates with the distance.
How to fix it
- Find the affected pages in Glimana. Open Issues › Near-duplicate pages. This is a site-level check, so the single row carries the full evidence: the URLs, values or days involved. The task under Tasks shows the same list and the expected gain.
-
Apply the fix on your platform.
- Group the affected pages by pattern (same template, same section).
- Decide per group: merge into one URL with 301s; keep one as canonical and point the others to it; or rewrite so each page covers a distinct intent.
- Prevent the pattern: block parameter variants in robots or canonicals, make tag archives noindex when they duplicate categories, and avoid "find-and-replace city name" pages unless each has unique local content.
Noindex tag, author and date archives in the SEO plugin if they duplicate category pages; disable attachment pages; set canonicals on filtered views.
Product variants and
/collections/x/products/yduplicates canonicalise to/products/yby default; keep the theme's canonical. Review apps that create landing pages per keyword.Canonicalise parameter variants; render one URL per piece of content.
-
Publish and clear caches. Save and publish, then make sure the crawler will see the new version: WordPress — clear the page cache in your cache plugin and purge the CDN; Shopify — theme and content changes go live on save, but purge any CDN in front of the store; custom code — deploy and purge the edge cache. Check in a private window or with
curl -s https://your-domain/page | grep -i '<title\|canonical\|robots'that the live HTML has changed; Glimana reads what the server sends, not what your browser has cached. - Verify in Glimana. Open the task under Tasks and click Mark as done. The affected pages are queued for a verification crawl within a few hours, and the task closes when the affected pages are gone, redirected, canonicalised or no longer similar. If the check still fails, the task returns to New with a note; when it passes, the task is listed under Resolved technical issues on Impact reports. To check sooner, use Re-crawl affected pages on the issue.
How the fix is verified
The next crawl re-computes the groups; the task closes when the affected pages are gone, redirected, canonicalised or no longer similar.