A 1.9% net gain hid 419,450 changed relationships
Between the Common Crawl releases cc-main-2025-oct-nov-dec and cc-main-2026-mar-apr-may, GitHub's referring-domain total went from 736,858 to 750,824. That is a gain of 13,966, or 1.9%. On its own, that number tells a client almost nothing.
The breakdown is where it gets interesting. 216,708 referring domains appeared that weren't in the previous release. 202,742 dropped off. 534,116 were present in both. The two moving buckets add up to 419,450 changed relationships, or 44% of the domains observed across the two releases. All of that churn is invisible if you only look at the total, because the gains and losses nearly cancel out.
Why churn is worth reporting on its own
A flat total with high churn tells you the client's backlink profile turned over substantially and still came out slightly ahead. Each of the three buckets is a different kind of work.
The 202,742 lost domains are a review list: some are dead links worth trying to reclaim, most are not worth the time, and you can only tell which is which by looking. The 216,708 new domains are the trail left by whatever happened last quarter, whether that's a campaign, syndication, or scraper noise. The 534,116 persistent domains are the stable base to keep an eye on. You get none of these breakdowns from the headline count.
What one referring domain becomes at page level
A domain-level graph says github.io links to github.com. Tier 2 enrichment turns that single edge into exact source pages, exact target URLs, link types, nofollow flags, and a basic editorial/noise classification from archived HTML.
For github.io alone, Linkwake fetched 7,214 archived HTML candidates from Common Crawl and found 17,678 GitHub link occurrences across 3,950 distinct source pages. We repeated the same uncapped enrichment for gitbook.io, webflow.io, and readthedocs.io to compare dense documentation links with broader hosted-site platforms.
The result changes the client conversation. A lost or gained domain is only a triage queue; exact source pages and target URLs tell you what to inspect. It also shows why raw source-page volume is not enough: webflow.io had more HTML candidates than github.io, but far fewer observed GitHub link occurrences in this crawl.
| Referrer | HTML fetched | Source pages | Occurrences | Target URLs | Editorial / noise / technical | Top target |
|---|---|---|---|---|---|---|
| github.io | 7,214 | 3,950 | 17,678 | 9,459 | 16,234 / 1,345 / 99 | github.com/rust-lang/rust/issues/88674 |
| gitbook.io | 649 | 169 | 5,687 | 3,515 | 5,680 / 7 / 0 | github.com/0xProject/0x-mesh/releases |
| webflow.io | 8,192 | 237 | 407 | 102 | 346 / 52 / 9 | github.com/ |
| readthedocs.io | 79 | 75 | 174 | 59 | 167 / 7 / 0 | github.com/scverse/anndata |
linked to github.com/040code/blog-oidc-github-actions-aws
linked to github.com/0xProject/0x-mesh/releases
linked to github.com/wonderunit/font-thicccboi
linked to github.com/scverse/anndata
Competitor gaps, with the usual caveat
GitHub dwarfs GitLab and Bitbucket on raw scale, so this isn't a like-for-like fight. The gap query is still useful. In the latest release, 14,967 domains linked to GitLab but not GitHub, and 3,881 linked to Bitbucket but not GitHub.
Most of those gap domains are junk: parked pages, spam TLDs, throwaway subdomains. That is true of any raw web-graph gap list, and we'd rather say so than pretend otherwise. The point isn't that the raw list is a prospect list. It's that you start from a graph-scale set you can filter down, instead of a blank page before a client call.
| Domain | Referring domains | Shared with GitHub | Gap vs GitHub |
|---|---|---|---|
| github.com | 750,824 | — | — |
| gitlab.com | 59,515 | 44,548 | 14,967 |
| bitbucket.org | 22,177 | 18,296 | 3,881 |
How this was measured
Every number here came from the live Linkwake API, backed by a ClickHouse store of Common Crawl's domain-level web graph. Each release covers roughly 100 to 180 million domains and 4 to 6 billion domain-to-domain links. A referring domain is a domain with at least one page linking to the target. Counts are domain-level, not page-level, so we're counting distinct linking sites, not individual links.
Common Crawl releases are periodic open-web snapshots, published roughly every few months, not daily rank data. That cadence is what makes them useful for portfolio monitoring: each release is a fresh movement report across your clients, competitors, and prospects, without suite pricing to watch referring-domain changes.