feat(seo-data): C2 — cannibalisation from Google's own data, one param away
The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.
CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.
fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
impressions, strongest page first inside each. Same auth, same quota family,
no new scope. `capped` reports a full row window rather than presenting a
truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.
Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.
Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.
30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.
The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.
Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
This commit is contained in:
@@ -340,8 +340,37 @@ When STEP 0/STEP 1 recorded a GSC account+property (not "none"):
|
||||
```bash
|
||||
bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query
|
||||
bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/"
|
||||
bash ~/.claude/lib/seo-data/fetch.sh cannibal --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90
|
||||
```
|
||||
|
||||
**`cannibal` — keyword cannibalisation, from Google's own data (C2).** Groups
|
||||
90 days of `query`+`page` rows and returns every query where 2+ of OUR pages
|
||||
compete, ranked by total impressions. The API always allowed multiple
|
||||
dimensions; this system only ever asked for one, so the conflict was invisible.
|
||||
|
||||
Read it:
|
||||
- `conflicts[]` → for each, the strongest page (most impressions) is listed
|
||||
first. That is usually the one to KEEP; the others either consolidate into
|
||||
it (301 + merge content) or get differentiated. Never "fix" this by deleting
|
||||
a page that has clicks — say what competes and let the user choose.
|
||||
- A conflict with a large impression total and every page beyond position 10
|
||||
is the real prize: Google can't decide which page to rank, so none rank.
|
||||
- `capped: true` → the row window was full; there are conflicts past the cut.
|
||||
Say so in §14 rather than presenting the list as exhaustive.
|
||||
- `status: degraded` → no GSC account. Cannibalisation is then **not
|
||||
auditable** — no substitute exists on-site. §14 line, do not guess it from
|
||||
title similarity.
|
||||
|
||||
**This is NOT the 30/70 rule, and do not merge the two.** Cannibalisation is
|
||||
a SERP fact Google measured. The 30/70 duplication rule is a content-similarity
|
||||
question with **no data source here**: measuring it properly needs main-content
|
||||
extraction (strip nav/header/footer), and without that a naive comparison of
|
||||
two same-template pages returns ~95% similar for every site, which is a
|
||||
confident false positive. So 30/70 stays an explicit LLM judgement over the
|
||||
≥3 same-family pages STEP 5 now samples for it — label it as judgement in the
|
||||
report, never as a measurement, and never quote a similarity percentage you
|
||||
did not compute.
|
||||
|
||||
Report: top queries; flag **QUICK WINS** = rows with position between 4
|
||||
and 10 AND high impressions (candidates to push onto page 1 with a
|
||||
title/meta/content tweak). Report index coverage from `inspect`. All
|
||||
|
||||
Reference in New Issue
Block a user