From 3a15643c2c5e68c22a1eb163b6257e5c815a44a7 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 11:55:00 +0200 Subject: [PATCH] =?UTF-8?q?feat(seo-data):=20C2=20=E2=80=94=20cannibalisat?= =?UTF-8?q?ion=20from=20Google's=20own=20data,=20one=20param=20away?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The inventory called this "no duplicate-content / cannibalisation detection". Splitting that into its two halves shows one is free and the other is a trap. CANNIBALISATION — free, and the data was already reachable. Search Analytics has always accepted several dimensions at once ("no limit to the number of dimensions that you can group by"); this engine only ever sent `"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So query+page — the pairing that exposes the conflict — was one parameter away and nobody asked. Same shape of win as W1. fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total impressions, strongest page first inside each. Same auth, same quota family, no new scope. `capped` reports a full row window rather than presenting a truncated list as exhaustive — same rule as COVERAGE and the sitemap caps. Grouping happens in the engine, deterministically: asking an LLM to group 1000 rows by query is arithmetic it should never be handed. Backward compatible: rows gained `keys` (the list the API actually returns); `key` stays as keys[0], so the single-dim quick-wins consumer is untouched. A test pins both. 30/70 DUPLICATION — deliberately NOT built, and this is the honest half. Measuring it needs main-content extraction (strip nav/header/footer). Without that, comparing two same-template pages returns ~95% similar for every site — a confident false positive, which is exactly the failure class the rest of this branch exists to remove. It stays an explicit LLM judgement over the >=3 same-family pages C1c now samples for it, labelled as judgement, never quoting a similarity percentage nobody computed. A wrong number would be worse than the current honest gap. The two must not be merged in the report either: cannibalisation is a SERP fact Google measured; 30/70 is a content question. The spec now says so. Verified: fixture with 3 pages on one query, 2 on another, 1 on a third → 2 conflicts, correct ranking, single-page query excluded; live dispatch degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite green; shellcheck + py_compile clean. --- agents/seo-analyzer.md | 29 ++++++++ lib/seo-data/README.md | 22 +++++++ lib/seo-data/fetch.sh | 4 +- .../fixtures-cannibal/gsc_queries.json | 7 ++ lib/seo-data/google_seo.py | 66 +++++++++++++++++-- lib/seo-data/seo-data.test.sh | 18 +++++ 6 files changed, 139 insertions(+), 7 deletions(-) create mode 100644 lib/seo-data/fixtures-cannibal/gsc_queries.json diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 6cf88ba..ae4843f 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -340,8 +340,37 @@ When STEP 0/STEP 1 recorded a GSC account+property (not "none"): ```bash bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/" +bash ~/.claude/lib/seo-data/fetch.sh cannibal --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 ``` +**`cannibal` — keyword cannibalisation, from Google's own data (C2).** Groups +90 days of `query`+`page` rows and returns every query where 2+ of OUR pages +compete, ranked by total impressions. The API always allowed multiple +dimensions; this system only ever asked for one, so the conflict was invisible. + +Read it: +- `conflicts[]` → for each, the strongest page (most impressions) is listed + first. That is usually the one to KEEP; the others either consolidate into + it (301 + merge content) or get differentiated. Never "fix" this by deleting + a page that has clicks — say what competes and let the user choose. +- A conflict with a large impression total and every page beyond position 10 + is the real prize: Google can't decide which page to rank, so none rank. +- `capped: true` → the row window was full; there are conflicts past the cut. + Say so in §14 rather than presenting the list as exhaustive. +- `status: degraded` → no GSC account. Cannibalisation is then **not + auditable** — no substitute exists on-site. §14 line, do not guess it from + title similarity. + +**This is NOT the 30/70 rule, and do not merge the two.** Cannibalisation is +a SERP fact Google measured. The 30/70 duplication rule is a content-similarity +question with **no data source here**: measuring it properly needs main-content +extraction (strip nav/header/footer), and without that a naive comparison of +two same-template pages returns ~95% similar for every site, which is a +confident false positive. So 30/70 stays an explicit LLM judgement over the +≥3 same-family pages STEP 5 now samples for it — label it as judgement in the +report, never as a measurement, and never quote a similarity percentage you +did not compute. + Report: top queries; flag **QUICK WINS** = rows with position between 4 and 10 AND high impressions (candidates to push onto page 1 with a title/meta/content tweak). Report index coverage from `inspect`. All diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index faf1792..60574f2 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -98,6 +98,28 @@ fetch.sh inspect --account client-a --property … --url https://ex.com/page • errors/warnings count issue INSTANCES; issues[] is deduped — the same issueMessage repeats across every affected item. +fetch.sh cannibal --account client-a --property … [--days 90] [--rows 1000] + → {"status":"ok","source":"gsc","days":90,"rows_scanned":1000,"capped":true, + "conflict_count":12, + "conflicts":[{"query":"plombier paris","pages":3,"total_impressions":2400, + "urls":[{"url":…,"clicks":…,"impressions":…,"position":…}]}]} + → {"status":"degraded","reason":"…"} # no account → NOT auditable + + Keyword cannibalisation from Google's own data: queries where 2+ of OUR + pages compete. Groups query+page rows; conflicts ranked by total + impressions, and within each the strongest page first. `capped:true` means + the row window was full — more conflicts exist past the cut, say so. + Same auth, same quota family, no new scope: the API always accepted several + dimensions at once, this engine only ever asked for one. + • NOT the 30/70 duplication rule. This is a SERP fact Google measured. + 30/70 is content similarity, which has no data source here — doing it + naively (compare two same-template pages without stripping nav/footer) + returns ~95% similar for every site, a confident false positive. It stays + an LLM judgement, labelled as one. + • `queries` now takes `--dim query,page` (comma-separated) and `--rows`. + Rows gained a `keys` list; `key` stays as keys[0], so the single-dim + consumer is untouched. + fetch.sh sitemap --url https://ex.com/sitemap.xml → {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0, "urls":["https://ex.com/", …]} diff --git a/lib/seo-data/fetch.sh b/lib/seo-data/fetch.sh index 8ca859d..65a7c33 100644 --- a/lib/seo-data/fetch.sh +++ b/lib/seo-data/fetch.sh @@ -27,7 +27,7 @@ _label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit cmd="${1:-}"; shift || true case "$cmd" in accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;; - crux|queries|inspect) + crux|queries|inspect|cannibal) exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;; # No auth, no Google: stdlib-only, runs even without the venv. sitemap) @@ -44,6 +44,6 @@ case "$cmd" in fi echo '{"status":"error","reason":"usage: fetch.sh forget {--label