feat(seo-data): C2 — cannibalisation from Google's own data, one param away

The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.

CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.

  fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
  impressions, strongest page first inside each. Same auth, same quota family,
  no new scope. `capped` reports a full row window rather than presenting a
  truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.

Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.

Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.

30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.

The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.

Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 11:55:00 +02:00
parent 04ccc5ad9b
commit 3a15643c2c
6 changed files with 139 additions and 7 deletions
+29
View File
@@ -340,8 +340,37 @@ When STEP 0/STEP 1 recorded a GSC account+property (not "none"):
```bash ```bash
bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query
bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/" bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/"
bash ~/.claude/lib/seo-data/fetch.sh cannibal --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90
``` ```
**`cannibal` — keyword cannibalisation, from Google's own data (C2).** Groups
90 days of `query`+`page` rows and returns every query where 2+ of OUR pages
compete, ranked by total impressions. The API always allowed multiple
dimensions; this system only ever asked for one, so the conflict was invisible.
Read it:
- `conflicts[]` → for each, the strongest page (most impressions) is listed
first. That is usually the one to KEEP; the others either consolidate into
it (301 + merge content) or get differentiated. Never "fix" this by deleting
a page that has clicks — say what competes and let the user choose.
- A conflict with a large impression total and every page beyond position 10
is the real prize: Google can't decide which page to rank, so none rank.
- `capped: true` → the row window was full; there are conflicts past the cut.
Say so in §14 rather than presenting the list as exhaustive.
- `status: degraded` → no GSC account. Cannibalisation is then **not
auditable** — no substitute exists on-site. §14 line, do not guess it from
title similarity.
**This is NOT the 30/70 rule, and do not merge the two.** Cannibalisation is
a SERP fact Google measured. The 30/70 duplication rule is a content-similarity
question with **no data source here**: measuring it properly needs main-content
extraction (strip nav/header/footer), and without that a naive comparison of
two same-template pages returns ~95% similar for every site, which is a
confident false positive. So 30/70 stays an explicit LLM judgement over the
≥3 same-family pages STEP 5 now samples for it — label it as judgement in the
report, never as a measurement, and never quote a similarity percentage you
did not compute.
Report: top queries; flag **QUICK WINS** = rows with position between 4 Report: top queries; flag **QUICK WINS** = rows with position between 4
and 10 AND high impressions (candidates to push onto page 1 with a and 10 AND high impressions (candidates to push onto page 1 with a
title/meta/content tweak). Report index coverage from `inspect`. All title/meta/content tweak). Report index coverage from `inspect`. All
+22
View File
@@ -98,6 +98,28 @@ fetch.sh inspect --account client-a --property … --url https://ex.com/page
• errors/warnings count issue INSTANCES; issues[] is deduped — the same • errors/warnings count issue INSTANCES; issues[] is deduped — the same
issueMessage repeats across every affected item. issueMessage repeats across every affected item.
fetch.sh cannibal --account client-a --property … [--days 90] [--rows 1000]
→ {"status":"ok","source":"gsc","days":90,"rows_scanned":1000,"capped":true,
"conflict_count":12,
"conflicts":[{"query":"plombier paris","pages":3,"total_impressions":2400,
"urls":[{"url":…,"clicks":…,"impressions":…,"position":…}]}]}
→ {"status":"degraded","reason":"…"} # no account → NOT auditable
Keyword cannibalisation from Google's own data: queries where 2+ of OUR
pages compete. Groups query+page rows; conflicts ranked by total
impressions, and within each the strongest page first. `capped:true` means
the row window was full — more conflicts exist past the cut, say so.
Same auth, same quota family, no new scope: the API always accepted several
dimensions at once, this engine only ever asked for one.
• NOT the 30/70 duplication rule. This is a SERP fact Google measured.
30/70 is content similarity, which has no data source here — doing it
naively (compare two same-template pages without stripping nav/footer)
returns ~95% similar for every site, a confident false positive. It stays
an LLM judgement, labelled as one.
• `queries` now takes `--dim query,page` (comma-separated) and `--rows`.
Rows gained a `keys` list; `key` stays as keys[0], so the single-dim
consumer is untouched.
fetch.sh sitemap --url https://ex.com/sitemap.xml fetch.sh sitemap --url https://ex.com/sitemap.xml
→ {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0, → {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0,
"urls":["https://ex.com/", …]} "urls":["https://ex.com/", …]}
+2 -2
View File
@@ -27,7 +27,7 @@ _label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit
cmd="${1:-}"; shift || true cmd="${1:-}"; shift || true
case "$cmd" in case "$cmd" in
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;; accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
crux|queries|inspect) crux|queries|inspect|cannibal)
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;; exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
# No auth, no Google: stdlib-only, runs even without the venv. # No auth, no Google: stdlib-only, runs even without the venv.
sitemap) sitemap)
@@ -44,6 +44,6 @@ case "$cmd" in
fi fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}' echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;; exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|sitemap|forget} [flags]"}' *) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|forget} [flags]"}'
exit 2 ;; exit 2 ;;
esac esac
@@ -0,0 +1,7 @@
{"rows":[
{"keys":["plombier paris","https://ex.com/plombier"],"clicks":40,"impressions":900,"ctr":0.044,"position":6.3},
{"keys":["plombier paris","https://ex.com/services/plomberie"],"clicks":3,"impressions":300,"ctr":0.010,"position":14.1},
{"keys":["urgence fuite","https://ex.com/urgence"],"clicks":5,"impressions":1200,"ctr":0.004,"position":8.9},
{"keys":["urgence fuite","https://ex.com/blog/fuite-que-faire"],"clicks":2,"impressions":800,"ctr":0.003,"position":11.4},
{"keys":["urgence fuite","https://ex.com/services/depannage"],"clicks":1,"impressions":400,"ctr":0.002,"position":19.2},
{"keys":["devis plomberie","https://ex.com/devis"],"clicks":9,"impressions":150,"ctr":0.060,"position":4.1}]}
+61 -5
View File
@@ -89,13 +89,16 @@ def _gsc_session(store_path, account):
return AuthorizedSession(creds) return AuthorizedSession(creds)
def _norm_queries(raw, dim): def _norm_queries(raw, dim):
# `keys` is the list the API actually returns (one entry per requested
# dimension); `key` stays as keys[0] so the single-dim consumer that reads
# it keeps working. Additive — nothing to migrate.
return {"status": "ok", "source": "gsc", "dimension": dim, "rows": [ return {"status": "ok", "source": "gsc", "dimension": dim, "rows": [
{"key": r["keys"][0], "clicks": r.get("clicks", 0), {"key": r["keys"][0], "keys": r["keys"], "clicks": r.get("clicks", 0),
"impressions": r.get("impressions", 0), "ctr": r.get("ctr", 0), "impressions": r.get("impressions", 0), "ctr": r.get("ctr", 0),
"position": r.get("position")} "position": r.get("position")}
for r in raw.get("rows", [])]} for r in raw.get("rows", [])]}
def queries(store_path, account, property, days=90, dim="query"): def queries(store_path, account, property, days=90, dim="query", rows=100):
raw = _mock("gsc_queries.json") raw = _mock("gsc_queries.json")
if raw is None: if raw is None:
sess = _gsc_session(store_path, account) sess = _gsc_session(store_path, account)
@@ -106,8 +109,12 @@ def queries(store_path, account, property, days=90, dim="query"):
import urllib.parse import urllib.parse
url = ("https://searchconsole.googleapis.com/webmasters/v3/sites/" url = ("https://searchconsole.googleapis.com/webmasters/v3/sites/"
+ urllib.parse.quote(property, safe="") + "/searchAnalytics/query") + urllib.parse.quote(property, safe="") + "/searchAnalytics/query")
# dim accepts a comma-separated list: the API groups by several
# dimensions at once ("no limit... but you cannot group by the same
# dimension twice"), and query+page is what exposes cannibalisation.
dims = [d.strip() for d in dim.split(",") if d.strip()]
r = sess.post(url, json={"startDate": start.isoformat(), "endDate": end.isoformat(), r = sess.post(url, json={"startDate": start.isoformat(), "endDate": end.isoformat(),
"dimensions": [dim], "rowLimit": 100}, timeout=30) "dimensions": dims, "rowLimit": rows}, timeout=30)
if r.status_code == 429: if r.status_code == 429:
return {"status": "degraded", "reason": "rate_limited"} return {"status": "degraded", "reason": "rate_limited"}
r.raise_for_status() r.raise_for_status()
@@ -147,6 +154,44 @@ def _norm_rich(ir):
"errors": errors, "warnings": warnings, "issues": msgs}) "errors": errors, "warnings": warnings, "issues": msgs})
return {"verdict": rr.get("verdict"), "types": types} return {"verdict": rr.get("verdict"), "types": types}
def _group_by_query(rows):
"""query+page rows -> {query: [row, …]}. Deterministic aggregation, not
judgement: the agent must not be asked to group 1000 rows by eye."""
by_q = {}
for r in rows:
keys = r.get("keys") or []
if len(keys) < 2:
continue
by_q.setdefault(keys[0], []).append(
{"url": keys[1], "clicks": r["clicks"],
"impressions": r["impressions"], "position": r["position"]})
return by_q
def cannibal(store_path, account, property, days=90, rows=1000):
"""Queries where 2+ of our own pages compete for the same term.
Google's own data says it; nothing in this system asked. Cannibalisation
is a SERP fact, not a content-similarity guess — do not confuse it with
the 30/70 duplication rule, which has no data source here."""
res = queries(store_path, account, property, days, "query,page", rows)
if res.get("status") != "ok":
return res
conflicts = []
for q, pages in _group_by_query(res["rows"]).items():
if len(pages) < 2:
continue
pages.sort(key=lambda p: p["impressions"], reverse=True)
conflicts.append({"query": q, "pages": len(pages),
"total_impressions": sum(p["impressions"] for p in pages),
"urls": pages})
conflicts.sort(key=lambda c: c["total_impressions"], reverse=True)
return {"status": "ok", "source": "gsc", "days": days,
"rows_scanned": len(res["rows"]),
# rows_scanned == rows means the window was FULL: there may be more
# conflicts past the cut. Reported, never silently truncated.
"capped": len(res["rows"]) >= rows,
"conflict_count": len(conflicts), "conflicts": conflicts}
def inspect(store_path, account, property, url): def inspect(store_path, account, property, url):
raw = _mock("gsc_inspect.json") raw = _mock("gsc_inspect.json")
if raw is None: if raw is None:
@@ -182,7 +227,15 @@ def _cli():
pq.add_argument("--account", required=True) pq.add_argument("--account", required=True)
pq.add_argument("--property", required=True) pq.add_argument("--property", required=True)
pq.add_argument("--days", type=int, default=90) pq.add_argument("--days", type=int, default=90)
pq.add_argument("--dim", default="query") pq.add_argument("--dim", default="query",
help="one dimension, or a comma-separated list (query,page)")
pq.add_argument("--rows", type=int, default=100)
pn = sub.add_parser("cannibal")
pn.add_argument("--store", required=True)
pn.add_argument("--account", required=True)
pn.add_argument("--property", required=True)
pn.add_argument("--days", type=int, default=90)
pn.add_argument("--rows", type=int, default=1000)
pi = sub.add_parser("inspect") pi = sub.add_parser("inspect")
pi.add_argument("--store", required=True) pi.add_argument("--store", required=True)
pi.add_argument("--account", required=True) pi.add_argument("--account", required=True)
@@ -193,7 +246,10 @@ def _cli():
print(json.dumps(crux(args.url, args.strategy), indent=2)) print(json.dumps(crux(args.url, args.strategy), indent=2))
elif args.cmd == "queries": elif args.cmd == "queries":
print(json.dumps(queries(args.store, args.account, args.property, print(json.dumps(queries(args.store, args.account, args.property,
args.days, args.dim), indent=2)) args.days, args.dim, args.rows), indent=2))
elif args.cmd == "cannibal":
print(json.dumps(cannibal(args.store, args.account, args.property,
args.days, args.rows), indent=2))
elif args.cmd == "inspect": elif args.cmd == "inspect":
print(json.dumps(inspect(args.store, args.account, args.property, print(json.dumps(inspect(args.store, args.account, args.property,
args.url), indent=2)) args.url), indent=2))
+18
View File
@@ -81,6 +81,24 @@ has "gsc degrades w/o creds" "$DEG" '"status": "degraded"'
has "gsc degrade reason" "$DEG" 'no_credentials' has "gsc degrade reason" "$DEG" 'no_credentials'
rm -rf "$TMP2" rm -rf "$TMP2"
echo "── cannibalisation ──"
# `keys` is additive: the single-dim consumer that reads `key` must not break
has "queries keeps key (compat)" "$Q" '"key": "plombier paris"'
has "queries adds keys list" "$Q" '"keys"'
CAN="$(SEO_DATA_MOCK_DIR="$SD/fixtures-cannibal" python3 "$SD/google_seo.py" cannibal \
--store "$S2" --account client-a --property sc-domain:ex.com)"
has "cannibal ok" "$CAN" '"status": "ok"'
# fixture: 3 pages on "urgence fuite", 2 on "plombier paris", 1 on "devis"
has "cannibal finds 2 conflicts" "$CAN" '"conflict_count": 2'
has "cannibal counts pages" "$CAN" '"pages": 3'
has "cannibal sums impressions" "$CAN" '"total_impressions": 2400'
hasnt "single-page query is not a conflict" "$CAN" 'devis plomberie'
# biggest conflict first, and inside it the strongest page first
CAN_FIRST="$(printf '%s' "$CAN" | python3 -c 'import sys,json; d=json.load(sys.stdin); print(d["conflicts"][0]["query"], d["conflicts"][0]["urls"][0]["url"])')"
check_first() { [ "$1" = "$2" ] && ok "$3" || no "$3" "got[$1]"; }
check_first "$CAN_FIRST" "urgence fuite https://ex.com/urgence" "cannibal ranks by impact"
has "cannibal reports the cap" "$CAN" '"capped": false'
echo "── sitemap ──" echo "── sitemap ──"
SM="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/sitemap.py" --url https://ex.com/sitemap.xml)" SM="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/sitemap.py" --url https://ex.com/sitemap.xml)"
has "sitemap ok" "$SM" '"status": "ok"' has "sitemap ok" "$SM" '"status": "ok"'