Files
claude_mac/lib/seo-data/README.md
T
Bastien Chanot a6d423b940 feat(seo-data): W1 — surface rich_results, the data inspect already threw away
google_seo.py:129 read only indexStatusResult out of the URL Inspection
response and discarded the rest. richResultsResult was already on the wire:
same call, same OAuth scope (webmasters.readonly), same quota. Google's own
structured-data verdict on the live indexed URL was being downloaded and
binned.

Plan correction: the TODO said "richresults verb". Wrong — a new verb means
a second POST to the same endpoint for a payload already received, on a
per-site quota, and nobody wants rich results without index status. Extended
inspect() instead; fetch.sh unchanged, no new verb, no new scope.

Design driven by the published schema, not by guesswork — two details I
would have got wrong:
- richResultsResult is OMITTED when Google detects none ("absent if none
  found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a
  caller cannot tell an absent key from a check that never ran. ABSENT means
  "none detected", never "invalid". The KeyError path is the real risk here,
  so it has its own fixture dir (fixtures-norich/) and its own tests.
- PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It
  never emits it now, and a test asserts the absence.

issues[] deduped (the same issueMessage repeats across every affected item),
errors/warnings count instances — scale from the counter, cause from the
message.

seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD
validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a
verified property, so its reach is the STEP 9 COVERAGE ratio, not the site.
Replacing a fake validator with a fake coverage promise would be no better.

This is what beats claude-seo: their README's "dual validator (Rich Results
Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of
their .py finds zero calls. This is Google's verdict, via auth already held.

Note: the new dedupe assertion trips SC2015 (A && B || C), same as the
pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for
house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh
shellcheck glob anyway.

Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end
and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean.
2026-07-16 20:51:30 +02:00

225 lines
11 KiB
Markdown

# seo-data — GSC + CrUX data layer for `/seo` FULL audits
Small, isolated engine that gives the `/seo` skill real Google data instead of
guesses: **Search Console** (queries, positions, indexation) and **CrUX**
(Core Web Vitals *field* data — real users, not lab simulation). It knows
nothing about SEO scoring; it only turns Google APIs into normalized JSON.
The `seo-analyzer` agent consumes that JSON in STEP 4 (Core Web Vitals) and
the new "Performance GSC" subsection; the `/seo` skill selects the account
and property in STEP 0 of a FULL audit (not needed for LOCAL).
Multi-account by design: the token store is keyed by a user-chosen label, and
every call takes `--account`/`--property` explicitly. Two audits running at
the same time (two sites, two sessions) never share mutable state — nothing
is written to disk during an audit, only at `make seo-connect`.
## Setup
One-time per Google account:
```bash
make seo-connect # from the claude-config repo
bash ~/.claude/lib/seo-data/connect.sh --label <label> # from ANY directory (venv must exist)
```
`make seo-connect` creates `~/.claude/.venv-seo-data/` (isolated venv, deps
pinned in `requirements.txt`), installs `google-auth`,
`google-auth-oauthlib`, `requests`, then delegates to `connect.sh`. The
wrapper sources `~/.claude/.env` internally, prefers the venv python, and
runs `connect.py`: it opens a browser for OAuth consent and takes a
**label** (e.g. `client-a`) to key the account — pick a name, not an email,
since the store never stores or requests the account's email. Once the venv
exists, `connect.sh` alone connects further accounts from anywhere (the
`/seo connect [label]` skill verb uses exactly this path).
Before running it, set these 3 keys in `~/.claude/.env` (the canonical
vault; `link.sh` only symlinks the repo's `.env` to it and warns with a
`cp .env.example .env` hint if it's missing — it never creates the vault
itself):
```bash
GOOGLE_OAUTH_CLIENT_ID=<your-client-id>.apps.googleusercontent.com
GOOGLE_OAUTH_CLIENT_SECRET=<your-client-secret>
CRUX_API_KEY=<your-crux-api-key>
```
- `GOOGLE_OAUTH_CLIENT_ID` / `GOOGLE_OAUTH_CLIENT_SECRET` — OAuth2 "Desktop
app" credentials from the Google Cloud Console (APIs & Services →
Credentials). Shared across every account you connect; the OAuth scope
requested is `https://www.googleapis.com/auth/webmasters.readonly` only
— read-only Search Console, nothing can be modified or deleted via this
token.
- `CRUX_API_KEY` — a Chrome UX Report API key (restrict it to CrUX +
PageSpeed in the Console). Get one at
https://developer.chrome.com/docs/crux/api. No OAuth involved: CrUX is
public field data, gated by API key only, independent of any connected
account.
`make seo-connect` is idempotent and rerunnable — connecting a second
account just runs it again with a different label; reusing an existing
label prompts to overwrite.
## `fetch.sh` contract
`lib/seo-data/fetch.sh` is the one stable entrypoint analyzers call. It
sources `~/.claude/.env`, prefers the isolated venv (falls back to system
`python3` for stdlib-only paths), dispatches to `google_seo.py` or
`tokenstore.py`, and never prints a secret to stdout or stderr.
```bash
fetch.sh accounts
→ {"status":"ok","accounts":[{"label":"…","properties":[…],"granted_at":"…"}]} # [] if none connected
fetch.sh crux --url https://ex.com [--strategy mobile|desktop]
→ {"status":"ok","source":"crux","lcp_p75_ms":…,"inp_p75_ms":…,"cls_p75":…} # a missing metric omits its key
→ {"status":"degraded","reason":"no_crux_key"|"no_field_data"|"rate_limited"}
# a 404 on page-level data retries at origin-level before degrading
fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--dim query|page]
→ {"status":"ok","source":"gsc","dimension":"query","rows":[{"key":"…","clicks":…,"impressions":…,"ctr":…,"position":…}]}
→ {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"}
fetch.sh inspect --account client-a --property … --url https://ex.com/page
→ {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…",
"rich_results":{"verdict":"PASS|FAIL|NEUTRAL|VERDICT_UNSPECIFIED|ABSENT",
"types":[{"type":"FAQ","items":2,"errors":2,"warnings":1,
"issues":["Missing field 'acceptedAnswer'"]}]}}
→ {"status":"degraded","reason":"…"}
rich_results rides the SAME URL-Inspection response — Google already sends
it, `inspect` used to discard it. No extra call, quota or OAuth scope.
It is the only programmatic structured-data validation in the system.
• verdict PARTIAL is never emitted — the API reserves it as unused.
• verdict ABSENT is SYNTHETIC (not a Google enum): the API omits
richResultsResult entirely when it detects no rich results. Surfaced
as a value rather than a missing key, because a caller cannot tell an
absent key apart from a check that never ran. ABSENT = "none
detected", never "invalid".
• errors/warnings count issue INSTANCES; issues[] is deduped — the same
issueMessage repeats across every affected item.
fetch.sh forget --label client-a
→ {"status":"ok","removed":true|false} # false = label wasn't in the store
fetch.sh forget --all
→ {"status":"ok","cleared":<n>} # n = accounts removed
```
Rules that hold for every subcommand:
- **JSON always on stdout, never empty.** Even an unexpected error (HTTP
403/5xx, timeout, DNS failure) prints
`{"status":"degraded","reason":"unexpected_error"}` — never a raw
traceback.
- **`status` is `"ok"` or `"degraded"` on exit 0; `"error"` on exit 2.**
Analyzers branch on this field; `"error"` only shows up on bad usage,
`reason` is informational otherwise.
- **Exit code 0 on `ok` and on `degraded`.** The engine never fails the
process just because Google data isn't available — that's a normal,
expected outcome the analyzer handles by falling back. **Exit code 2**
is reserved for bad usage: unknown subcommand, missing required flag,
invalid argument — those paths emit `{"status":"error",...}` instead.
- **`--store` is accepted uniformly** by every subcommand for consistent
`fetch.sh` dispatch, even though `crux` ignores it (CrUX needs no
account).
- **Never prints a secret.** No env var, refresh token, or access token
ever reaches stdout or stderr, including in error paths.
Two env vars exist for testing, never for normal use:
`SEO_DATA_ENV_FILE` overrides which env file is sourced (tests point it at
`/dev/null` so a real `~/.claude/.env` on the machine can never leak into a
test run), and `SEO_DATA_DEBUG=1` re-enables stderr for local debugging
(stderr is suppressed by default so library warnings can't leak a secret
into an agent's context).
## Token store
`~/.claude/seo-data/tokens.json` — refresh tokens, keyed by the label chosen
at `make seo-connect`, one entry per connected account:
```json
{
"version": 1,
"accounts": {
"client-a": {
"refresh_token": "<opaque>",
"scopes": ["https://www.googleapis.com/auth/webmasters.readonly"],
"granted_at": "2026-07-09T12:00:00+00:00",
"properties": ["sc-domain:site-a.com", "https://www.site-a.com/"]
}
}
}
```
Security posture:
- **File `0600`, directory `0700`.** `tokenstore.save_account` re-asserts
both permissions on every write.
- **Written only at `connect` time, atomically.** `tmp` → `fsync` →
`os.replace` (atomic rename), under an exclusive `fcntl` lock, so two
simultaneous `make seo-connect` runs can't corrupt the file. Audits never
write to this file — access tokens are exchanged in memory and never
persisted, so two audits running concurrently never contend on it.
- **Keyed by label, not email.** Identifying accounts by email would
require widening the OAuth scope just for identification; the label the
user picks at connect time is sufficient and keeps the scope at
`webmasters.readonly` only (least privilege).
- **Refresh tokens are redacted from `list`.** `fetch.sh accounts` (and
`tokenstore.py list`) return label, properties, and `granted_at` only —
the `refresh_token` field is intentionally never included in that output.
- **Allowlisted in gitleaks.** The store lives under `~/.claude/`, outside
this repo, so it's never committed directly — but `make scan-secrets`
also sweeps `~/.claude` for stray copies of secrets. `.gitleaks.toml` has
an explicit `[allowlist].paths` entry for
`(^|/)\.claude/seo-data/tokens\.json$`, the same treatment
`~/.claude/.env` already gets, so a legitimate local secret store doesn't
drown real findings in false positives.
- **Also gitignored** (`.venv-seo-data/` and `seo-data/tokens.json` in
`.gitignore`) as a second, belt-and-suspenders guard in case a relative
path ever put either under the repo tree.
- **Removal is local-only.** `fetch.sh forget --label <x>` / `--all` (the
`/seo forget` skill verb) deletes the stored refresh token — it does NOT
revoke the OAuth grant at Google's end. For a real revocation, visit
https://myaccount.google.com/permissions with the account concerned and
remove the app's access; the deleted local token then becomes useless
everywhere, including to anyone who copied it beforehand.
## Graceful degradation
Missing API key, no connected account, or a revoked/expired token is a
**normal outcome, not a failure**:
- No `CRUX_API_KEY` → `crux` returns `{"status":"degraded","reason":"no_crux_key"}`.
- No account connected, or the store has no refresh token for the given
`--account` → `queries`/`inspect` return
`{"status":"degraded","reason":"no_credentials"}`.
- Refresh token revoked at Google's end → `{"status":"degraded","reason":"token_revoked"}`
(a transient network blip during refresh is classified
`"network_error"` instead, so a flaky connection never forces the user
back through OAuth).
- Rate limited (HTTP 429) on any Google API → `{"status":"degraded","reason":"rate_limited"}`.
In every case: **exit code 0**, valid JSON on stdout, no crash. The `/seo`
FULL audit continues on the anonymous PageSpeed API (lab data) instead of
CrUX field data, and the report surfaces the fix as a user action:
`make seo-connect`. `doctor.sh` also flags both non-fatally as `WARN`: a
missing `CRUX_API_KEY` warns on its own, while no connected Google account
is the one that names `make seo-connect`.
## Testing
```bash
make test
# or, to run only this engine's suite:
bash lib/seo-data/seo-data.test.sh
```
The suite is network-free: `google_seo.py` reads fixtures from
`lib/seo-data/fixtures/` (`crux_mobile.json`, `gsc_queries.json`,
`gsc_inspect.json`) whenever `SEO_DATA_MOCK_DIR` is set, instead of calling
Google's APIs. Degradation paths run with real env vars unset (`env -u
CRUX_API_KEY`, `env -u SEO_DATA_MOCK_DIR`) to exercise the no-key/no-creds
branches deterministically. Every `fetch.sh` invocation in the tests also
sets `SEO_DATA_ENV_FILE=/dev/null` so a machine with a live
`~/.claude/.env` never lets real credentials leak into a test run.