Files
claude_mac/lib/seo-data/README.md
T
Bastien Chanot a6d423b940 feat(seo-data): W1 — surface rich_results, the data inspect already threw away
google_seo.py:129 read only indexStatusResult out of the URL Inspection
response and discarded the rest. richResultsResult was already on the wire:
same call, same OAuth scope (webmasters.readonly), same quota. Google's own
structured-data verdict on the live indexed URL was being downloaded and
binned.

Plan correction: the TODO said "richresults verb". Wrong — a new verb means
a second POST to the same endpoint for a payload already received, on a
per-site quota, and nobody wants rich results without index status. Extended
inspect() instead; fetch.sh unchanged, no new verb, no new scope.

Design driven by the published schema, not by guesswork — two details I
would have got wrong:
- richResultsResult is OMITTED when Google detects none ("absent if none
  found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a
  caller cannot tell an absent key from a check that never ran. ABSENT means
  "none detected", never "invalid". The KeyError path is the real risk here,
  so it has its own fixture dir (fixtures-norich/) and its own tests.
- PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It
  never emits it now, and a test asserts the absence.

issues[] deduped (the same issueMessage repeats across every affected item),
errors/warnings count instances — scale from the counter, cause from the
message.

seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD
validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a
verified property, so its reach is the STEP 9 COVERAGE ratio, not the site.
Replacing a fake validator with a fake coverage promise would be no better.

This is what beats claude-seo: their README's "dual validator (Rich Results
Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of
their .py finds zero calls. This is Google's verdict, via auth already held.

Note: the new dedupe assertion trips SC2015 (A && B || C), same as the
pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for
house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh
shellcheck glob anyway.

Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end
and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean.
2026-07-16 20:51:30 +02:00

11 KiB

seo-data — GSC + CrUX data layer for /seo FULL audits

Small, isolated engine that gives the /seo skill real Google data instead of guesses: Search Console (queries, positions, indexation) and CrUX (Core Web Vitals field data — real users, not lab simulation). It knows nothing about SEO scoring; it only turns Google APIs into normalized JSON. The seo-analyzer agent consumes that JSON in STEP 4 (Core Web Vitals) and the new "Performance GSC" subsection; the /seo skill selects the account and property in STEP 0 of a FULL audit (not needed for LOCAL).

Multi-account by design: the token store is keyed by a user-chosen label, and every call takes --account/--property explicitly. Two audits running at the same time (two sites, two sessions) never share mutable state — nothing is written to disk during an audit, only at make seo-connect.

Setup

One-time per Google account:

make seo-connect                                        # from the claude-config repo
bash ~/.claude/lib/seo-data/connect.sh --label <label>  # from ANY directory (venv must exist)

make seo-connect creates ~/.claude/.venv-seo-data/ (isolated venv, deps pinned in requirements.txt), installs google-auth, google-auth-oauthlib, requests, then delegates to connect.sh. The wrapper sources ~/.claude/.env internally, prefers the venv python, and runs connect.py: it opens a browser for OAuth consent and takes a label (e.g. client-a) to key the account — pick a name, not an email, since the store never stores or requests the account's email. Once the venv exists, connect.sh alone connects further accounts from anywhere (the /seo connect [label] skill verb uses exactly this path).

Before running it, set these 3 keys in ~/.claude/.env (the canonical vault; link.sh only symlinks the repo's .env to it and warns with a cp .env.example .env hint if it's missing — it never creates the vault itself):

GOOGLE_OAUTH_CLIENT_ID=<your-client-id>.apps.googleusercontent.com
GOOGLE_OAUTH_CLIENT_SECRET=<your-client-secret>
CRUX_API_KEY=<your-crux-api-key>
  • GOOGLE_OAUTH_CLIENT_ID / GOOGLE_OAUTH_CLIENT_SECRET — OAuth2 "Desktop app" credentials from the Google Cloud Console (APIs & Services → Credentials). Shared across every account you connect; the OAuth scope requested is https://www.googleapis.com/auth/webmasters.readonly only — read-only Search Console, nothing can be modified or deleted via this token.
  • CRUX_API_KEY — a Chrome UX Report API key (restrict it to CrUX + PageSpeed in the Console). Get one at https://developer.chrome.com/docs/crux/api. No OAuth involved: CrUX is public field data, gated by API key only, independent of any connected account.

make seo-connect is idempotent and rerunnable — connecting a second account just runs it again with a different label; reusing an existing label prompts to overwrite.

fetch.sh contract

lib/seo-data/fetch.sh is the one stable entrypoint analyzers call. It sources ~/.claude/.env, prefers the isolated venv (falls back to system python3 for stdlib-only paths), dispatches to google_seo.py or tokenstore.py, and never prints a secret to stdout or stderr.

fetch.sh accounts
  → {"status":"ok","accounts":[{"label":"…","properties":[…],"granted_at":"…"}]}   # [] if none connected

fetch.sh crux    --url https://ex.com [--strategy mobile|desktop]
  → {"status":"ok","source":"crux","lcp_p75_ms":…,"inp_p75_ms":…,"cls_p75":…}      # a missing metric omits its key
  → {"status":"degraded","reason":"no_crux_key"|"no_field_data"|"rate_limited"}
  # a 404 on page-level data retries at origin-level before degrading

fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--dim query|page]
  → {"status":"ok","source":"gsc","dimension":"query","rows":[{"key":"…","clicks":…,"impressions":…,"ctr":…,"position":…}]}
  → {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"}

fetch.sh inspect --account client-a --property … --url https://ex.com/page
  → {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…",
     "rich_results":{"verdict":"PASS|FAIL|NEUTRAL|VERDICT_UNSPECIFIED|ABSENT",
                     "types":[{"type":"FAQ","items":2,"errors":2,"warnings":1,
                               "issues":["Missing field 'acceptedAnswer'"]}]}}
  → {"status":"degraded","reason":"…"}

  rich_results rides the SAME URL-Inspection response — Google already sends
  it, `inspect` used to discard it. No extra call, quota or OAuth scope.
  It is the only programmatic structured-data validation in the system.
    • verdict PARTIAL is never emitted — the API reserves it as unused.
    • verdict ABSENT is SYNTHETIC (not a Google enum): the API omits
      richResultsResult entirely when it detects no rich results. Surfaced
      as a value rather than a missing key, because a caller cannot tell an
      absent key apart from a check that never ran. ABSENT = "none
      detected", never "invalid".
    • errors/warnings count issue INSTANCES; issues[] is deduped — the same
      issueMessage repeats across every affected item.

fetch.sh forget --label client-a
  → {"status":"ok","removed":true|false}          # false = label wasn't in the store

fetch.sh forget --all
  → {"status":"ok","cleared":<n>}                 # n = accounts removed

Rules that hold for every subcommand:

  • JSON always on stdout, never empty. Even an unexpected error (HTTP 403/5xx, timeout, DNS failure) prints {"status":"degraded","reason":"unexpected_error"} — never a raw traceback.
  • status is "ok" or "degraded" on exit 0; "error" on exit 2. Analyzers branch on this field; "error" only shows up on bad usage, reason is informational otherwise.
  • Exit code 0 on ok and on degraded. The engine never fails the process just because Google data isn't available — that's a normal, expected outcome the analyzer handles by falling back. Exit code 2 is reserved for bad usage: unknown subcommand, missing required flag, invalid argument — those paths emit {"status":"error",...} instead.
  • --store is accepted uniformly by every subcommand for consistent fetch.sh dispatch, even though crux ignores it (CrUX needs no account).
  • Never prints a secret. No env var, refresh token, or access token ever reaches stdout or stderr, including in error paths.

Two env vars exist for testing, never for normal use: SEO_DATA_ENV_FILE overrides which env file is sourced (tests point it at /dev/null so a real ~/.claude/.env on the machine can never leak into a test run), and SEO_DATA_DEBUG=1 re-enables stderr for local debugging (stderr is suppressed by default so library warnings can't leak a secret into an agent's context).

Token store

~/.claude/seo-data/tokens.json — refresh tokens, keyed by the label chosen at make seo-connect, one entry per connected account:

{
  "version": 1,
  "accounts": {
    "client-a": {
      "refresh_token": "<opaque>",
      "scopes": ["https://www.googleapis.com/auth/webmasters.readonly"],
      "granted_at": "2026-07-09T12:00:00+00:00",
      "properties": ["sc-domain:site-a.com", "https://www.site-a.com/"]
    }
  }
}

Security posture:

  • File 0600, directory 0700. tokenstore.save_account re-asserts both permissions on every write.
  • Written only at connect time, atomically. tmp → fsync → os.replace (atomic rename), under an exclusive fcntl lock, so two simultaneous make seo-connect runs can't corrupt the file. Audits never write to this file — access tokens are exchanged in memory and never persisted, so two audits running concurrently never contend on it.
  • Keyed by label, not email. Identifying accounts by email would require widening the OAuth scope just for identification; the label the user picks at connect time is sufficient and keeps the scope at webmasters.readonly only (least privilege).
  • Refresh tokens are redacted from list. fetch.sh accounts (and tokenstore.py list) return label, properties, and granted_at only — the refresh_token field is intentionally never included in that output.
  • Allowlisted in gitleaks. The store lives under ~/.claude/, outside this repo, so it's never committed directly — but make scan-secrets also sweeps ~/.claude for stray copies of secrets. .gitleaks.toml has an explicit [allowlist].paths entry for (^|/)\.claude/seo-data/tokens\.json$, the same treatment ~/.claude/.env already gets, so a legitimate local secret store doesn't drown real findings in false positives.
  • Also gitignored (.venv-seo-data/ and seo-data/tokens.json in .gitignore) as a second, belt-and-suspenders guard in case a relative path ever put either under the repo tree.
  • Removal is local-only. fetch.sh forget --label <x> / --all (the /seo forget skill verb) deletes the stored refresh token — it does NOT revoke the OAuth grant at Google's end. For a real revocation, visit https://myaccount.google.com/permissions with the account concerned and remove the app's access; the deleted local token then becomes useless everywhere, including to anyone who copied it beforehand.

Graceful degradation

Missing API key, no connected account, or a revoked/expired token is a normal outcome, not a failure:

  • No CRUX_API_KEY → crux returns {"status":"degraded","reason":"no_crux_key"}.
  • No account connected, or the store has no refresh token for the given --account → queries/inspect return {"status":"degraded","reason":"no_credentials"}.
  • Refresh token revoked at Google's end → {"status":"degraded","reason":"token_revoked"} (a transient network blip during refresh is classified "network_error" instead, so a flaky connection never forces the user back through OAuth).
  • Rate limited (HTTP 429) on any Google API → {"status":"degraded","reason":"rate_limited"}.

In every case: exit code 0, valid JSON on stdout, no crash. The /seo FULL audit continues on the anonymous PageSpeed API (lab data) instead of CrUX field data, and the report surfaces the fix as a user action: make seo-connect. doctor.sh also flags both non-fatally as WARN: a missing CRUX_API_KEY warns on its own, while no connected Google account is the one that names make seo-connect.

Testing

make test
# or, to run only this engine's suite:
bash lib/seo-data/seo-data.test.sh

The suite is network-free: google_seo.py reads fixtures from lib/seo-data/fixtures/ (crux_mobile.json, gsc_queries.json, gsc_inspect.json) whenever SEO_DATA_MOCK_DIR is set, instead of calling Google's APIs. Degradation paths run with real env vars unset (env -u CRUX_API_KEY, env -u SEO_DATA_MOCK_DIR) to exercise the no-key/no-creds branches deterministically. Every fetch.sh invocation in the tests also sets SEO_DATA_ENV_FILE=/dev/null so a machine with a live ~/.claude/.env never lets real credentials leak into a test run.