fix(seo): B1 KILLED — Common Crawl backlinks measured, not assumed

The plan said Common Crawl was the free backlink source and the 70/100 cap
was therefore mandatory. Measured before building, and both premises die.

HEAD against data.commoncrawl.org, live:
  cc-main-2026-feb-mar-apr-domain-edges.txt.gz    17.3 GB   gzipped
  cc-main-2026-feb-mar-apr-domain-ranks.txt.gz     2.3 GB
  cc-main-2026-feb-mar-apr-domain-vertices.txt.gz  879 MB

Finding one domain's inbound links means scanning the edges file end to end,
per audit. That is not slow, it is non-viable — and abusive toward a
nonprofit serving the data free.

Worse, the reference implementation everyone points at
(claude-seo scripts/commoncrawl_graph.py:169) does this:

    max_compressed_bytes = 500 * 1024 * 1024   # 500 MiB safety cap
    if total_downloaded > max_compressed_bytes: break

500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source
ID — so it reads an arbitrary slice of source domains and reports whatever
backlinks happened to be in it, as a backlink profile, capped at "70/100
health". Nothing in the output says 3%. That is a random sample wearing a
measurement's clothes: the exact failure class this branch exists to remove,
and I was one step from copying it.

B2 dies with B1: nothing left to cap.

CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand
mentions only, backlinks + authority declared unauditable in §14 — is the
FINAL state, not a placeholder waiting for data. Corrected my own I1 text,
which pointed at Common Crawl as the "nearest free source": that sends a
future reader into a 17 GB dead end. The §14 line now records what was
measured and why no number beats a fabricated one.

Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free
source") was wrong for the same reason. The only free viable backlink source
is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on
the client's Bing account. That raises W2's value; it does not unblock it.

Verified: full suite green, seo-data 144 pass / 0 fail.
This commit is contained in:
Bastien Chanot
2026-07-17 13:17:18 +02:00
parent 20d3082542
commit d6b8edc8ea
2 changed files with 54 additions and 14 deletions
+23 -6
View File
@@ -944,13 +944,30 @@ that reaches a client via `/client-handover`. A low mention count is a low
mention count — it is NOT evidence of a weak backlink profile.
Mandatory §14 line whenever depth=FULL, verbatim:
`Backlinks / domain authority — NOT audited: no backlink index wired.
Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs /
Semrush / Majestic. The Off-page score above prices in brand mentions only.`
`Backlinks / domain authority — NOT audited: no free backlink index is
practical, and none is wired. Commercial: Ahrefs / Semrush / Majestic. The
Off-page score above prices in brand mentions only.`
Weight deliberately unchanged despite the narrower scope: re-deriving it
now, then again when a backlink source lands, would churn historical
scores twice. Revisit the 10/15% only when the axis widens back.
**This is the final state, not a placeholder (B1 killed, 2026-07-17.)** The
free options were measured, not assumed:
- **GSC has no links endpoint.** The Search Console API exposes exactly
Search Analytics, Sitemaps, Sites, URL Inspection. The Links report is
UI-only.
- **Common Crawl's hyperlinkgraph is 17.3 GB gzipped** for the domain-edges
file alone (+879 MB vertices, +2.3 GB ranks), measured live. Finding one
domain's inbound links means scanning all of it, per audit. Not slow —
non-viable, and abusive toward a nonprofit serving it free. The reference
implementation everyone cites caps its download at 500 MiB, i.e. **2.9% of
the edges file**, and reports whatever that arbitrary slice contained as a
backlink profile. That is a random sample wearing a measurement's clothes,
which is precisely what this axis note exists to prevent.
- **Bing Webmaster's `GetUrlLinks` is the only free, viable source** — but it
is first-party only (your verified properties), so it can never cover a
competitor, and it needs the client's Bing account. See W2, deferred.
So: no number here beats a fabricated one. Weight deliberately unchanged —
re-deriving it for an axis that is not going to widen would churn historical
scores for nothing.
### LOCAL depth — 4 axes