Files
claude/lib/seo-data/fetch.sh
T
Bastien Chanot 20d3082542 feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright
Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
2026-07-17 12:44:43 +02:00

54 lines
2.5 KiB
Bash

#!/usr/bin/env bash
# Stable entrypoint for the seo-data engine. JSON on stdout; exit 0 on ok/degrade,
# exit 2 on bad usage. Never prints secrets.
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
ENV_FILE="${SEO_DATA_ENV_FILE:-${HOME}/.claude/.env}" # canonical; tests override to /dev/null
STORE="${SEO_DATA_STORE:-${HOME}/.claude/seo-data/tokens.json}"
VENV_PY="${HOME}/.claude/.venv-seo-data/bin/python3"
# Library stderr must never leak a secret into agent context — suppress it
# globally unless explicitly debugging (SEO_DATA_DEBUG=1 restores it).
[ -n "${SEO_DATA_DEBUG:-}" ] || exec 2>/dev/null
# Load secrets quietly (sourced, never echoed).
if [ -f "$ENV_FILE" ]; then
set -a; # shellcheck source=/dev/null
. "$ENV_FILE"; set +a
fi
# Prefer the isolated venv (has google-auth); fall back to system python3 for
# stdlib-only paths (accounts / mock / degrade).
PY="python3"; [ -x "$VENV_PY" ] && PY="$VENV_PY"
# Whole-string label guard (shell-safe ASCII). POSIX `case` in a C-locale
# subshell — newline-proof and locale-independent, unlike a per-line grep.
_label_safe() ( LC_ALL=C; case "$1" in ''|[!A-Za-z0-9]*|*[!A-Za-z0-9._-]*) exit 1;; esac )
cmd="${1:-}"; shift || true
case "$cmd" in
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
crux|queries|inspect|cannibal)
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
# No auth, no Google: stdlib-only, runs even without the venv.
sitemap)
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
rendercheck)
exec "$PY" "$HERE/render_check.py" --store "$STORE" "$@" ;;
linkgraph)
exec "$PY" "$HERE/linkgraph.py" --store "$STORE" "$@" ;;
forget)
# forget --label <label> → drop one account; forget --all → empty the store.
# Local removal only — does NOT revoke the grant at Google's end.
# Label charset guard: store keys stay shell-safe wherever an agent
# interpolates them into a command line (defense-in-depth vs injection).
if [ "${1:-}" = "--all" ]; then
exec "$PY" "$HERE/tokenstore.py" clear --file "$STORE"
elif [ "${1:-}" = "--label" ] && _label_safe "${2:-}"; then
exec "$PY" "$HERE/tokenstore.py" remove --file "$STORE" --label "$2"
fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|forget} [flags]"}'
exit 2 ;;
esac