forked from bchanot/claude
The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.
CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.
fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
impressions, strongest page first inside each. Same auth, same quota family,
no new scope. `capped` reports a full row window rather than presenting a
truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.
Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.
Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.
30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.
The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.
Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
50 lines
2.3 KiB
Bash
50 lines
2.3 KiB
Bash
#!/usr/bin/env bash
|
|
# Stable entrypoint for the seo-data engine. JSON on stdout; exit 0 on ok/degrade,
|
|
# exit 2 on bad usage. Never prints secrets.
|
|
set -uo pipefail
|
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
|
ENV_FILE="${SEO_DATA_ENV_FILE:-${HOME}/.claude/.env}" # canonical; tests override to /dev/null
|
|
STORE="${SEO_DATA_STORE:-${HOME}/.claude/seo-data/tokens.json}"
|
|
VENV_PY="${HOME}/.claude/.venv-seo-data/bin/python3"
|
|
|
|
# Library stderr must never leak a secret into agent context — suppress it
|
|
# globally unless explicitly debugging (SEO_DATA_DEBUG=1 restores it).
|
|
[ -n "${SEO_DATA_DEBUG:-}" ] || exec 2>/dev/null
|
|
|
|
# Load secrets quietly (sourced, never echoed).
|
|
if [ -f "$ENV_FILE" ]; then
|
|
set -a; # shellcheck source=/dev/null
|
|
. "$ENV_FILE"; set +a
|
|
fi
|
|
# Prefer the isolated venv (has google-auth); fall back to system python3 for
|
|
# stdlib-only paths (accounts / mock / degrade).
|
|
PY="python3"; [ -x "$VENV_PY" ] && PY="$VENV_PY"
|
|
|
|
# Whole-string label guard (shell-safe ASCII). POSIX `case` in a C-locale
|
|
# subshell — newline-proof and locale-independent, unlike a per-line grep.
|
|
_label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit 1;; esac )
|
|
|
|
cmd="${1:-}"; shift || true
|
|
case "$cmd" in
|
|
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
|
crux|queries|inspect|cannibal)
|
|
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
|
# No auth, no Google: stdlib-only, runs even without the venv.
|
|
sitemap)
|
|
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
|
|
forget)
|
|
# forget --label <label> → drop one account; forget --all → empty the store.
|
|
# Local removal only — does NOT revoke the grant at Google's end.
|
|
# Label charset guard: store keys stay shell-safe wherever an agent
|
|
# interpolates them into a command line (defense-in-depth vs injection).
|
|
if [ "${1:-}" = "--all" ]; then
|
|
exec "$PY" "$HERE/tokenstore.py" clear --file "$STORE"
|
|
elif [ "${1:-}" = "--label" ] && _label_safe "${2:-}"; then
|
|
exec "$PY" "$HERE/tokenstore.py" remove --file "$STORE" --label "$2"
|
|
fi
|
|
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
|
exit 2 ;;
|
|
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|forget} [flags]"}'
|
|
exit 2 ;;
|
|
esac
|