forked from bchanot/claude
feat(seo-data): C2 — cannibalisation from Google's own data, one param away
The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.
CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.
fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
impressions, strongest page first inside each. Same auth, same quota family,
no new scope. `capped` reports a full row window rather than presenting a
truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.
Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.
Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.
30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.
The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.
Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
This commit is contained in:
@@ -27,7 +27,7 @@ _label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit
|
||||
cmd="${1:-}"; shift || true
|
||||
case "$cmd" in
|
||||
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
||||
crux|queries|inspect)
|
||||
crux|queries|inspect|cannibal)
|
||||
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
||||
# No auth, no Google: stdlib-only, runs even without the venv.
|
||||
sitemap)
|
||||
@@ -44,6 +44,6 @@ case "$cmd" in
|
||||
fi
|
||||
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
||||
exit 2 ;;
|
||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|sitemap|forget} [flags]"}'
|
||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|forget} [flags]"}'
|
||||
exit 2 ;;
|
||||
esac
|
||||
|
||||
Reference in New Issue
Block a user