From d6b8edc8eab36a1d53b09a25665206223c04fb89 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 13:17:18 +0200 Subject: [PATCH] =?UTF-8?q?fix(seo):=20B1=20KILLED=20=E2=80=94=20Common=20?= =?UTF-8?q?Crawl=20backlinks=20measured,=20not=20assumed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The plan said Common Crawl was the free backlink source and the 70/100 cap was therefore mandatory. Measured before building, and both premises die. HEAD against data.commoncrawl.org, live: cc-main-2026-feb-mar-apr-domain-edges.txt.gz 17.3 GB gzipped cc-main-2026-feb-mar-apr-domain-ranks.txt.gz 2.3 GB cc-main-2026-feb-mar-apr-domain-vertices.txt.gz 879 MB Finding one domain's inbound links means scanning the edges file end to end, per audit. That is not slow, it is non-viable — and abusive toward a nonprofit serving the data free. Worse, the reference implementation everyone points at (claude-seo scripts/commoncrawl_graph.py:169) does this: max_compressed_bytes = 500 * 1024 * 1024 # 500 MiB safety cap if total_downloaded > max_compressed_bytes: break 500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source ID — so it reads an arbitrary slice of source domains and reports whatever backlinks happened to be in it, as a backlink profile, capped at "70/100 health". Nothing in the output says 3%. That is a random sample wearing a measurement's clothes: the exact failure class this branch exists to remove, and I was one step from copying it. B2 dies with B1: nothing left to cap. CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand mentions only, backlinks + authority declared unauditable in §14 — is the FINAL state, not a placeholder waiting for data. Corrected my own I1 text, which pointed at Common Crawl as the "nearest free source": that sends a future reader into a 17 GB dead end. The §14 line now records what was measured and why no number beats a fabricated one. Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free source") was wrong for the same reason. The only free viable backlink source is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on the client's Bing account. That raises W2's value; it does not unblock it. Verified: full suite green, seo-data 144 pass / 0 fail. --- .claude/tasks/TODO.md | 39 +++++++++++++++++++++++++++++++-------- agents/seo-analyzer.md | 29 +++++++++++++++++++++++------ 2 files changed, 54 insertions(+), 14 deletions(-) diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index d60a28b..08e2efa 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -13,8 +13,10 @@ NEXT: H1 (SSRF/injection guard) → C1 (sitemap crawl). Human merge gate: all - **B3 KILLED** — GSC Links API does not exist. Verified against the API reference: Search Console v1 exposes exactly Search Analytics, Sitemaps, Sites, URL Inspection. A subagent hallucinated it; I doubted it in the - plan and the doubt was right. Common Crawl is the ONLY free backlink - source → the 70/100 cap is mandatory, not optional. + plan and the doubt was right. (Its follow-on — "so Common Crawl is the + only free source, and the 70/100 cap is mandatory" — was ALSO wrong: see + B1/B2 KILLED below. Common Crawl is a 17 GB dead end, and Bing's + GetUrlLinks is the only viable free source, first-party only.) - **I1 was an over-correction** — "Off-page has ZERO data" was overstated (relayed from a subagent, unverified). Brand mentions ARE gathered (STEP 6). Narrowed the axis definition instead of N/A-ing it; weights @@ -30,6 +32,26 @@ NEXT: H1 (SSRF/injection guard) → C1 (sitemap crawl). Human merge gate: all follow-up. Today only $DOMAIN (user-typed) is interpolated. After C1, N URLs from a REMOTE sitemap flow into shell commands and fetch targets. +### B1/B2 (Common Crawl backlinks) — KILLED 2026-07-17, measured not assumed +The plan said Common Crawl was the free backlink source and the 70/100 cap +was therefore mandatory. Both premises are dead: +- domain-edges.txt.gz = **17.3 GB gzipped** (+879 MB vertices, +2.3 GB + ranks), measured live via HEAD. Finding one domain's inbound links means + scanning all of it, per audit. Non-viable, and abusive toward a nonprofit. +- The implementation everyone cites (claude-seo commoncrawl_graph.py:169) + caps at `500 MiB` = **2.9% of the edges file**, and reports what that + arbitrary slice held as a backlink profile. A random sample presented as a + measurement — the exact failure class this branch exists to remove. We + nearly copied it. +- B2 dies with B1: nothing to cap. +CONSEQUENCE: I1's narrowed Off-page axis (brand mentions only, backlinks + +authority declared unauditable in §14) is the FINAL state, not a placeholder. +Its §14 line was corrected — it used to point at Common Crawl as "nearest +free source", which is a 17 GB dead end. +RAISES W2's VALUE: Bing's GetUrlLinks is now the ONLY free viable backlink +source. First-party only (never a competitor), still blocked on the client's +Bing account. + ### W2 (Bing) — DEFERRED, blocked on a real-world test Killed after 4 challenge rounds. User's model: client sites live on CLIENT Bing accounts, so a per-user API key means one key per client account. @@ -109,12 +131,13 @@ as NEW VERBS. No new architecture. - [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated as checks with no command to compute them. C1 unblocks real computation. -### AXE 3 — Off-page real (upgrades I1) -- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph - (data.commoncrawl.org/projects/hyperlinkgraph), free, no key. -- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap - health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine - exactly. +### AXE 3 — Off-page real (upgrades I1) — SUPERSEDED, see B1/B2 KILLED above +- [x] ~~B1 `backlinks` verb — Common Crawl hyperlinkgraph~~ KILLED: edges file + measured at 17.3 GB gzipped. Non-viable per audit; the reference impl + caps at 500 MiB = 2.9% of the graph and calls the remainder a backlink + profile. +- [x] ~~B2 Honest cap at 70/100~~ KILLED with B1: nothing left to cap. + I1's narrowed axis is the final state. - [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth already there" — I doubt it: Search Console API v3 has no links endpoint (links report is UI-only AFAIK). Verify before planning on it. Do not diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 6f315ff..6ec050d 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -944,13 +944,30 @@ that reaches a client via `/client-handover`. A low mention count is a low mention count — it is NOT evidence of a weak backlink profile. Mandatory §14 line whenever depth=FULL, verbatim: -`Backlinks / domain authority — NOT audited: no backlink index wired. -Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs / -Semrush / Majestic. The Off-page score above prices in brand mentions only.` +`Backlinks / domain authority — NOT audited: no free backlink index is +practical, and none is wired. Commercial: Ahrefs / Semrush / Majestic. The +Off-page score above prices in brand mentions only.` -Weight deliberately unchanged despite the narrower scope: re-deriving it -now, then again when a backlink source lands, would churn historical -scores twice. Revisit the 10/15% only when the axis widens back. +**This is the final state, not a placeholder (B1 killed, 2026-07-17.)** The +free options were measured, not assumed: +- **GSC has no links endpoint.** The Search Console API exposes exactly + Search Analytics, Sitemaps, Sites, URL Inspection. The Links report is + UI-only. +- **Common Crawl's hyperlinkgraph is 17.3 GB gzipped** for the domain-edges + file alone (+879 MB vertices, +2.3 GB ranks), measured live. Finding one + domain's inbound links means scanning all of it, per audit. Not slow — + non-viable, and abusive toward a nonprofit serving it free. The reference + implementation everyone cites caps its download at 500 MiB, i.e. **2.9% of + the edges file**, and reports whatever that arbitrary slice contained as a + backlink profile. That is a random sample wearing a measurement's clothes, + which is precisely what this axis note exists to prevent. +- **Bing Webmaster's `GetUrlLinks` is the only free, viable source** — but it + is first-party only (your verified properties), so it can never cover a + competitor, and it needs the client's Bing account. See W2, deferred. + +So: no number here beats a fabricated one. Weight deliberately unchanged — +re-deriving it for an axis that is not going to widen would churn historical +scores for nothing. ### LOCAL depth — 4 axes