Linkwake
Open web-graph data

Common Crawl backlink tool

Explore referring domains and link movement using the open Common Crawl web graph instead of a closed proprietary index.

The portfolio problem

Proprietary backlink indexes are powerful, but their collection methods and portfolio economics can be opaque. Common Crawl provides a transparent, independently available foundation for large-scale domain research.

Know the source

Every analysis is tied to a named Common Crawl web-graph release and comparison window.

Reproduce the method

The underlying dataset is public, so researchers can inspect the release rather than trusting an unexplained score.

Monitor at portfolio scale

Domain-level relationships make broad monitoring practical without pretending to replace page-level indexes.

A practical workflow

From raw movement to a useful decision

  1. 1

    Common Crawl samples the web

    Its crawlers fetch a large, changing portion of public pages and publish the resulting corpus.

  2. 2

    The web graph aggregates links

    Page links are reduced to relationships between registrable domains for each release.

  3. 3

    Linkwake computes movement

    Adjacent releases are compared to classify referring domains as new, lost, or persistent.

Worked example

What this looks like in practice

A relationship from publisher.example to client.example means at least one crawled page on the publisher linked to the client during that graph release; it does not mean every page or every known link was captured.

Coverage and limitations

Common Crawl is a sample, not a complete or real-time census. Coverage varies by release, robots rules, crawl priority, and site accessibility. Counts will differ from Google, Ahrefs, Majestic, Semrush, and Search Console.

Read the Common Crawl methodology or explore the underlying approach in our original research.

Questions, answered

Frequently asked questions

Is Common Crawl the same as Google's index?

No. It is an independent public crawl with different goals, coverage, scheduling, and processing.

Why can a known backlink be missing?

The source page may not have been crawled in that release, may block crawlers, may require JavaScript, or may fall outside the published graph.

What is the data good for?

It is especially useful for transparent research, broad domain-level comparisons, portfolio monitoring, and repeatable release-to-release analysis.

See Linkwake with real portfolio data

Explore the demo, check any domain for free, or start a five-domain workspace without a credit card.