The plan said Common Crawl was the free backlink source and the 70/100 cap
was therefore mandatory. Measured before building, and both premises die.
HEAD against data.commoncrawl.org, live:
cc-main-2026-feb-mar-apr-domain-edges.txt.gz 17.3 GB gzipped
cc-main-2026-feb-mar-apr-domain-ranks.txt.gz 2.3 GB
cc-main-2026-feb-mar-apr-domain-vertices.txt.gz 879 MB
Finding one domain's inbound links means scanning the edges file end to end,
per audit. That is not slow, it is non-viable — and abusive toward a
nonprofit serving the data free.
Worse, the reference implementation everyone points at
(claude-seo scripts/commoncrawl_graph.py:169) does this:
max_compressed_bytes = 500 * 1024 * 1024 # 500 MiB safety cap
if total_downloaded > max_compressed_bytes: break
500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source
ID — so it reads an arbitrary slice of source domains and reports whatever
backlinks happened to be in it, as a backlink profile, capped at "70/100
health". Nothing in the output says 3%. That is a random sample wearing a
measurement's clothes: the exact failure class this branch exists to remove,
and I was one step from copying it.
B2 dies with B1: nothing left to cap.
CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand
mentions only, backlinks + authority declared unauditable in §14 — is the
FINAL state, not a placeholder waiting for data. Corrected my own I1 text,
which pointed at Common Crawl as the "nearest free source": that sends a
future reader into a 17 GB dead end. The §14 line now records what was
measured and why no number beats a fabricated one.
Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free
source") was wrong for the same reason. The only free viable backlink source
is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on
the client's Bing account. That raises W2's value; it does not unblock it.
Verified: full suite green, seo-data 144 pass / 0 fail.