diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 6cf88ba..ae4843f 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -340,8 +340,37 @@ When STEP 0/STEP 1 recorded a GSC account+property (not "none"): ```bash bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/" +bash ~/.claude/lib/seo-data/fetch.sh cannibal --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 ``` +**`cannibal` — keyword cannibalisation, from Google's own data (C2).** Groups +90 days of `query`+`page` rows and returns every query where 2+ of OUR pages +compete, ranked by total impressions. The API always allowed multiple +dimensions; this system only ever asked for one, so the conflict was invisible. + +Read it: +- `conflicts[]` → for each, the strongest page (most impressions) is listed + first. That is usually the one to KEEP; the others either consolidate into + it (301 + merge content) or get differentiated. Never "fix" this by deleting + a page that has clicks — say what competes and let the user choose. +- A conflict with a large impression total and every page beyond position 10 + is the real prize: Google can't decide which page to rank, so none rank. +- `capped: true` → the row window was full; there are conflicts past the cut. + Say so in §14 rather than presenting the list as exhaustive. +- `status: degraded` → no GSC account. Cannibalisation is then **not + auditable** — no substitute exists on-site. §14 line, do not guess it from + title similarity. + +**This is NOT the 30/70 rule, and do not merge the two.** Cannibalisation is +a SERP fact Google measured. The 30/70 duplication rule is a content-similarity +question with **no data source here**: measuring it properly needs main-content +extraction (strip nav/header/footer), and without that a naive comparison of +two same-template pages returns ~95% similar for every site, which is a +confident false positive. So 30/70 stays an explicit LLM judgement over the +≥3 same-family pages STEP 5 now samples for it — label it as judgement in the +report, never as a measurement, and never quote a similarity percentage you +did not compute. + Report: top queries; flag **QUICK WINS** = rows with position between 4 and 10 AND high impressions (candidates to push onto page 1 with a title/meta/content tweak). Report index coverage from `inspect`. All diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index faf1792..60574f2 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -98,6 +98,28 @@ fetch.sh inspect --account client-a --property … --url https://ex.com/page • errors/warnings count issue INSTANCES; issues[] is deduped — the same issueMessage repeats across every affected item. +fetch.sh cannibal --account client-a --property … [--days 90] [--rows 1000] + → {"status":"ok","source":"gsc","days":90,"rows_scanned":1000,"capped":true, + "conflict_count":12, + "conflicts":[{"query":"plombier paris","pages":3,"total_impressions":2400, + "urls":[{"url":…,"clicks":…,"impressions":…,"position":…}]}]} + → {"status":"degraded","reason":"…"} # no account → NOT auditable + + Keyword cannibalisation from Google's own data: queries where 2+ of OUR + pages compete. Groups query+page rows; conflicts ranked by total + impressions, and within each the strongest page first. `capped:true` means + the row window was full — more conflicts exist past the cut, say so. + Same auth, same quota family, no new scope: the API always accepted several + dimensions at once, this engine only ever asked for one. + • NOT the 30/70 duplication rule. This is a SERP fact Google measured. + 30/70 is content similarity, which has no data source here — doing it + naively (compare two same-template pages without stripping nav/footer) + returns ~95% similar for every site, a confident false positive. It stays + an LLM judgement, labelled as one. + • `queries` now takes `--dim query,page` (comma-separated) and `--rows`. + Rows gained a `keys` list; `key` stays as keys[0], so the single-dim + consumer is untouched. + fetch.sh sitemap --url https://ex.com/sitemap.xml → {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0, "urls":["https://ex.com/", …]} diff --git a/lib/seo-data/fetch.sh b/lib/seo-data/fetch.sh index 8ca859d..65a7c33 100644 --- a/lib/seo-data/fetch.sh +++ b/lib/seo-data/fetch.sh @@ -27,7 +27,7 @@ _label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit cmd="${1:-}"; shift || true case "$cmd" in accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;; - crux|queries|inspect) + crux|queries|inspect|cannibal) exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;; # No auth, no Google: stdlib-only, runs even without the venv. sitemap) @@ -44,6 +44,6 @@ case "$cmd" in fi echo '{"status":"error","reason":"usage: fetch.sh forget {--label