Find pages that nearly repeat each other
Group pages whose text is estimated at least 80% similar, with the phrases they share, sampled Googlebot activity and honest coverage limits.
Crawl Audit now groups pages whose extracted text is nearly the same, alongside the exact-match check released earlier. Each group names a representative page, shows every saved member's estimated similarity to it and lists sample phrases they all share, so you can see why pages were grouped before opening them. Googlebot activity for the saved pages helps you decide what to review first.
The similarity is an estimate of how much text two pages share, not a search penalty score or an indexing verdict. Pages that differ only in numbers, dates or a product name score very high because they share almost all their text; the shared phrases show whether the overlap is template or content. Identical pages stay in the exact duplicate check and join a near-duplicate group as a single member.
Coverage explains which pages could be compared and which could not. Pages need about 45 words of extracted text, text without word spaces is not yet comparable, and groups keep up to 20 saved members with the true total shown separately. Start a new Salience crawl to collect the evidence; crawls captured before this release show the check as unavailable.
See Crawl Audit help for how to read the estimate and its limits.