forked from bchanot/claude
Merge feature/gsc-crux-data-layer into develop
This commit is contained in:
@@ -4,3 +4,12 @@
|
|||||||
# Used by: lib/toggle-external.sh enable|disable magic
|
# Used by: lib/toggle-external.sh enable|disable magic
|
||||||
# Get a key at: https://21st.dev/magic (dashboard → API keys)
|
# Get a key at: https://21st.dev/magic (dashboard → API keys)
|
||||||
MAGIC_API_KEY=your_21st_dev_magic_api_key_here
|
MAGIC_API_KEY=your_21st_dev_magic_api_key_here
|
||||||
|
|
||||||
|
# ── Google SEO data layer (lib/seo-data) — used by /seo FULL ──
|
||||||
|
# OAuth Desktop client: GCP console → APIs & Services → Credentials → OAuth client (Desktop).
|
||||||
|
# Scope requested at consent: webmasters.readonly. One-time setup: make seo-connect
|
||||||
|
GOOGLE_OAUTH_CLIENT_ID=<your-client-id.apps.googleusercontent.com>
|
||||||
|
GOOGLE_OAUTH_CLIENT_SECRET=<your-client-secret>
|
||||||
|
# CrUX + PageSpeed API key (GCP console → Credentials → API key, restricted to those APIs).
|
||||||
|
# Get it: https://developer.chrome.com/docs/crux/api
|
||||||
|
CRUX_API_KEY=<your-crux-api-key>
|
||||||
|
|||||||
@@ -113,6 +113,11 @@ install-*.log
|
|||||||
.env.*
|
.env.*
|
||||||
!.env.example
|
!.env.example
|
||||||
|
|
||||||
|
# seo-data engine local artifacts (live under ~/.claude, never committed)
|
||||||
|
.venv-seo-data/
|
||||||
|
seo-data/tokens.json
|
||||||
|
__pycache__/
|
||||||
|
|
||||||
# OS
|
# OS
|
||||||
.DS_Store
|
.DS_Store
|
||||||
Thumbs.db
|
Thumbs.db
|
||||||
|
|||||||
@@ -34,4 +34,7 @@ paths = [
|
|||||||
# for stray COPIES of secrets outside this file; flagging the vault
|
# for stray COPIES of secrets outside this file; flagging the vault
|
||||||
# itself on every run is pure noise, not signal.
|
# itself on every run is pure noise, not signal.
|
||||||
'''(^|/)\.env$''',
|
'''(^|/)\.env$''',
|
||||||
|
# seo-data OAuth token store — legitimate local secret (like ~/.claude/.env),
|
||||||
|
# 0600, outside git. Allowlisted so `make scan-secrets` doesn't flag the vault.
|
||||||
|
'''(^|/)\.claude/seo-data/tokens\.json$''',
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
.PHONY: help install plugin link doctor update new-skill profile profile-list profile-current profile-reset onboard test scan-secrets
|
.PHONY: help install plugin link doctor update new-skill profile profile-list profile-current profile-reset onboard test scan-secrets seo-connect
|
||||||
|
|
||||||
help: ## Show available commands
|
help: ## Show available commands
|
||||||
@grep -E '^[a-zA-Z_-]+:.*##' $(MAKEFILE_LIST) | awk 'BEGIN {FS = ":.*## "}; {printf " make %-14s %s\n", $$1, $$2}'
|
@grep -E '^[a-zA-Z_-]+:.*##' $(MAKEFILE_LIST) | awk 'BEGIN {FS = ":.*## "}; {printf " make %-14s %s\n", $$1, $$2}'
|
||||||
@@ -22,8 +22,14 @@ onboard: link ## Onboard an existing project (run from the project directory)
|
|||||||
@echo "Open Claude Code in your project directory and run: /onboard"
|
@echo "Open Claude Code in your project directory and run: /onboard"
|
||||||
@echo "Or with hints: /onboard Python FastAPI monorepo"
|
@echo "Or with hints: /onboard Python FastAPI monorepo"
|
||||||
|
|
||||||
|
seo-connect: ## Connect a Google account for /seo FULL (creates venv, OAuth consent)
|
||||||
|
@python3 -m venv "$$HOME/.claude/.venv-seo-data"
|
||||||
|
@"$$HOME/.claude/.venv-seo-data/bin/pip" install -q -r lib/seo-data/requirements.txt
|
||||||
|
@bash -c 'read -r -p "Label for this account (e.g. client-a): " label; \
|
||||||
|
"$$HOME/.claude/.venv-seo-data/bin/python3" lib/seo-data/connect.py --label "$$label"'
|
||||||
|
|
||||||
test: ## Run deterministic tests (lib/tests/*.test.sh + lib/gitflow-test.sh + lib/tests/run-*.sh)
|
test: ## Run deterministic tests (lib/tests/*.test.sh + lib/gitflow-test.sh + lib/tests/run-*.sh)
|
||||||
@fail=0; for t in lib/tests/*.test.sh lib/gitflow-test.sh lib/tests/run-*.sh; do \
|
@fail=0; for t in lib/tests/*.test.sh lib/seo-data/*.test.sh lib/gitflow-test.sh lib/tests/run-*.sh; do \
|
||||||
echo "== $$t"; \
|
echo "== $$t"; \
|
||||||
case "$$(basename "$$t")" in \
|
case "$$(basename "$$t")" in \
|
||||||
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || fail=1 ;; \
|
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || fail=1 ;; \
|
||||||
|
|||||||
@@ -69,6 +69,11 @@ Therefore: submit to GSC + Bing Webmaster minimum on every FULL audit.
|
|||||||
- **Google Search Console** (FREE) — https://search.google.com/search-console
|
- **Google Search Console** (FREE) — https://search.google.com/search-console
|
||||||
Covers Google search + AI Overviews grounding. URL inspection tool
|
Covers Google search + AI Overviews grounding. URL inspection tool
|
||||||
requests live re-indexing (faster than waiting for crawl).
|
requests live re-indexing (faster than waiting for crawl).
|
||||||
|
- **Connexion GSC pour /seo (données réelles)** — `make seo-connect`
|
||||||
|
(depuis le repo claude-config, une fois par compte) : consentement
|
||||||
|
OAuth lecture seule (webmasters.readonly), stocke un refresh token
|
||||||
|
local (0600). Ensuite /seo FULL lit requêtes/positions/indexation
|
||||||
|
sans réinvite.
|
||||||
- **IndexNow protocol** (FREE) — https://www.indexnow.org
|
- **IndexNow protocol** (FREE) — https://www.indexnow.org
|
||||||
Proactive ping to Bing + Yandex + Seznam + DuckDuckGo. One-line
|
Proactive ping to Bing + Yandex + Seznam + DuckDuckGo. One-line
|
||||||
API call per URL change. Plugins: Yoast (built-in), RankMath,
|
API call per URL change. Plugins: Yoast (built-in), RankMath,
|
||||||
|
|||||||
+47
-5
@@ -198,12 +198,18 @@ verify WebFetch + WebSearch available. If missing:
|
|||||||
|
|
||||||
```
|
```
|
||||||
PLUGIN CHECK
|
PLUGIN CHECK
|
||||||
curl/Bash : YES (always)
|
curl/Bash : YES (always)
|
||||||
WebFetch : YES / NO / N/A (LOCAL)
|
WebFetch : YES / NO / N/A (LOCAL)
|
||||||
WebSearch : YES / NO / N/A (LOCAL)
|
WebSearch : YES / NO / N/A (LOCAL)
|
||||||
STATUS : READY | DEGRADED (missing: <list>)
|
GSC/CrUX creds : READY (account: <label>) | DEGRADED (no account — anonymous PageSpeed only)
|
||||||
|
STATUS : READY | DEGRADED (missing: <list>)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
GSC/CrUX creds status comes from the `(account, property)` passed in
|
||||||
|
context (STEP 1). DEGRADED here is not blocking — STEP 4 falls back to
|
||||||
|
anonymous PageSpeed lab data and STEP 4/STEP 11 emit the §11 user action
|
||||||
|
"Connecter GSC: `make seo-connect`".
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## STEP 4 — LIVE TECHNICAL AUDIT `[FULL only]`
|
## STEP 4 — LIVE TECHNICAL AUDIT `[FULL only]`
|
||||||
@@ -244,7 +250,22 @@ Evaluate each present/missing:
|
|||||||
- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web
|
- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web
|
||||||
Vitals 2.0
|
Vitals 2.0
|
||||||
|
|
||||||
Use PageSpeed Insights API (no auth needed for basic usage):
|
When a GSC account+property were passed in context, fetch CrUX field
|
||||||
|
data first (**tilde path mandatory** — this agent runs from the
|
||||||
|
audited project's directory, not the claude-config repo):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh crux --url "https://$DOMAIN" --strategy mobile
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh crux --url "https://$DOMAIN" --strategy desktop
|
||||||
|
```
|
||||||
|
|
||||||
|
If `status=ok`, use `lcp_p75_ms` / `inp_p75_ms` / `cls_p75` as the
|
||||||
|
PRIMARY CWV figures (75th percentile, real users). Keep the PageSpeed
|
||||||
|
lab run below as a SECONDARY diagnostic. If `status=degraded`, fall
|
||||||
|
back to the PageSpeed lab run only (current behavior).
|
||||||
|
|
||||||
|
Use PageSpeed Insights API (no auth needed for basic usage) — SECONDARY
|
||||||
|
diagnostic, or PRIMARY when CrUX degraded:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -s "https://www.googleapis.com/pagespeedonline/v5/runPagespeed?url=https://$DOMAIN&strategy=mobile&category=PERFORMANCE&category=ACCESSIBILITY&category=BEST_PRACTICES&category=SEO" \
|
curl -s "https://www.googleapis.com/pagespeedonline/v5/runPagespeed?url=https://$DOMAIN&strategy=mobile&category=PERFORMANCE&category=ACCESSIBILITY&category=BEST_PRACTICES&category=SEO" \
|
||||||
@@ -257,6 +278,23 @@ Extract (via jq if available, otherwise WebFetch to transform):
|
|||||||
- `lighthouseResult.audits.cumulative-layout-shift.numericValue`
|
- `lighthouseResult.audits.cumulative-layout-shift.numericValue`
|
||||||
- Mobile + desktop separately
|
- Mobile + desktop separately
|
||||||
|
|
||||||
|
### Performance GSC (90 j) `[FULL only, account+property present]`
|
||||||
|
|
||||||
|
When STEP 0/STEP 1 recorded a GSC account+property (not "none"):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/"
|
||||||
|
```
|
||||||
|
|
||||||
|
Report: top queries; flag **QUICK WINS** = rows with position between 4
|
||||||
|
and 10 AND high impressions (candidates to push onto page 1 with a
|
||||||
|
title/meta/content tweak). Report index coverage from `inspect`. All
|
||||||
|
emitted into SEO.md §2 (technical) and §8 (quick wins).
|
||||||
|
|
||||||
|
If `status=degraded` → note it in §2 and emit the §11 user action
|
||||||
|
"Connecter GSC: `make seo-connect`".
|
||||||
|
|
||||||
### SEO technical files
|
### SEO technical files
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -573,6 +611,10 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
|||||||
| Competitive position | 5% | 10% | |
|
| Competitive position | 5% | 10% | |
|
||||||
| Legal compliance | 10% | 5% | |
|
| Legal compliance | 10% | 5% | |
|
||||||
|
|
||||||
|
**Technical axis note:** CWV scored on CrUX field data (75th percentile,
|
||||||
|
real users, from STEP 4) when available; otherwise lab PageSpeed
|
||||||
|
Lighthouse run.
|
||||||
|
|
||||||
### LOCAL depth — 4 axes
|
### LOCAL depth — 4 axes
|
||||||
|
|
||||||
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
||||||
|
|||||||
@@ -0,0 +1,964 @@
|
|||||||
|
# GSC + CrUX Data Layer — Implementation Plan
|
||||||
|
|
||||||
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||||
|
|
||||||
|
**Goal:** Give `/seo` (+`/geo`) FULL audits real Google data — Search Console queries/positions/indexation + CrUX field Core Web Vitals — via an isolated, secure, multi-account data engine.
|
||||||
|
|
||||||
|
**Architecture:** A self-contained engine under `lib/seo-data/` (bash entrypoint `fetch.sh` → Python helpers in an isolated venv) fetches GSC + CrUX and emits normalized JSON on stdout. The existing `seo-analyzer` agent consumes that JSON during FULL audits; the `/seo` dispatcher selects account+property in STEP 0. Secrets live in the `~/.claude/.env` vault (OAuth app + CrUX key) plus a label-keyed token store `~/.claude/seo-data/tokens.json`.
|
||||||
|
|
||||||
|
**Tech Stack:** Bash (entrypoint, tests), Python 3.14 (`google-auth`, `google-auth-oauthlib`, `requests` — pinned, in a dedicated venv), GSC Search Console API v3 + URL Inspection, CrUX API.
|
||||||
|
|
||||||
|
**Spec:** `docs/superpowers/specs/2026-07-09-gsc-crux-data-layer-design.md` (transient — delete after ship+doc+capitalize).
|
||||||
|
|
||||||
|
## Global Constraints
|
||||||
|
|
||||||
|
Every task's requirements implicitly include these (verbatim from the spec):
|
||||||
|
|
||||||
|
- **Security first.** Secrets never in git, files `0600` / dirs `0700`, OAuth scope **exactly** `https://www.googleapis.com/auth/webmasters.readonly`, no secret ever printed to stdout/stderr/report.
|
||||||
|
- **Offline-testable.** Third-party imports (`google.*`, `requests`) are **lazy** — imported only inside real OAuth/HTTP code paths. The `accounts`, mock (`SEO_DATA_MOCK_DIR` set), and degraded paths run on **stdlib only**, no venv, no network. `make test` never hits the network.
|
||||||
|
- **Graceful degradation (fail-open audit).** Missing creds / missing venv / revoked token / HTTP 429 → JSON `{"status":"degraded","reason":"…"}` on stdout with **exit 0**. Bad CLI usage → exit 2.
|
||||||
|
- **Multi-account, no shared state.** Account + property are **explicit arguments** on every `fetch.sh` call. No "current account" global. Store writes only happen during `connect` (atomic `tmp`→`fsync`→`rename` under `fcntl` lock); audits are read-only.
|
||||||
|
- **Store keyed by user label**, not email (keeps scope minimal). Properties discovered via `sites.list` (already in scope).
|
||||||
|
- **Canonical env path.** Read secrets from `~/.claude/.env` (canonical), never `$REPO/.env` (symlink may be absent on a fresh machine).
|
||||||
|
- **Repo test convention.** Bash test following the repo idiom (helpers `tf`/`tr_`/`tn` + `PASS`/`FAIL` counters, final line `[ "$FAIL" -eq 0 ]`). The engine test lives at `lib/seo-data/seo-data.test.sh` — co-located with the engine, **deliberately NOT under `lib/tests/`** which the `config-protection.sh` hook gates as a guardrail dir. Task 6 extends the `make test` target to also discover `lib/seo-data/*.test.sh`. During TDD, run it directly: `bash lib/seo-data/seo-data.test.sh`.
|
||||||
|
- **No commit attribution trailers** (no `Co-Authored-By`, no `Claude-Session`).
|
||||||
|
- **Branch:** all commits on `feature/gsc-crux-data-layer` (already created).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## File Structure
|
||||||
|
|
||||||
|
**Engine (created):**
|
||||||
|
- `lib/seo-data/tokenstore.py` — label-keyed token store I/O (atomic + locked). Stdlib only.
|
||||||
|
- `lib/seo-data/google_seo.py` — CrUX + GSC calls, OAuth refresh (lazy), mock mode, normalization → JSON.
|
||||||
|
- `lib/seo-data/connect.py` — one-time OAuth consent + `sites.list` discovery + persist to store.
|
||||||
|
- `lib/seo-data/fetch.sh` — bash entrypoint: source env, pick python, dispatch, degrade, redact.
|
||||||
|
- `lib/seo-data/requirements.txt` — pinned deps.
|
||||||
|
- `lib/seo-data/README.md` — usage contract.
|
||||||
|
|
||||||
|
**Tests (created):**
|
||||||
|
- `lib/seo-data/seo-data.test.sh` — deterministic bash test (drives CLIs against fixtures, checks locks).
|
||||||
|
- `lib/seo-data/fixtures/*.json` — synthetic API responses (no real secret/PII).
|
||||||
|
|
||||||
|
**Wiring (modified):**
|
||||||
|
- `.env.example`, `install.sh`, `Makefile`, `doctor.sh`, `.gitleaks.toml`, `.gitignore`.
|
||||||
|
|
||||||
|
**Integration (modified):**
|
||||||
|
- `agents/seo-analyzer.md`, `skills/seo/SKILL.md`, `agents/resources/automation-catalog.md`.
|
||||||
|
|
||||||
|
**Interface contract (used across tasks):**
|
||||||
|
```
|
||||||
|
tokenstore.py (module + CLI: python3 tokenstore.py {list|set} --file PATH …)
|
||||||
|
load(path) -> dict
|
||||||
|
list_accounts(path) -> list[dict] # [{label, properties, granted_at}] NO refresh_token
|
||||||
|
get_refresh_token(path, label) -> str | None
|
||||||
|
save_account(path, label, refresh_token, scopes: list[str], properties: list[str]) -> None
|
||||||
|
|
||||||
|
google_seo.py (module + CLI: python3 google_seo.py {crux|queries|inspect} …)
|
||||||
|
crux(url, strategy='mobile') -> dict
|
||||||
|
queries(store_path, account, property, days=90, dim='query') -> dict
|
||||||
|
inspect(store_path, account, property, url) -> dict
|
||||||
|
# all return {"status":"ok"|"degraded", ...}
|
||||||
|
|
||||||
|
fetch.sh {accounts|crux|queries|inspect} [flags] -> JSON on stdout
|
||||||
|
|
||||||
|
connect.py (CLI: python3 connect.py --label LABEL)
|
||||||
|
run_consent(client_id, client_secret, scopes) -> str # refresh_token
|
||||||
|
discover_properties(refresh_token, client_id, client_secret) -> list[str]
|
||||||
|
persist(store_path, label, refresh_token, scopes, properties) -> None
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 1: Token store (`tokenstore.py`)
|
||||||
|
|
||||||
|
Label-keyed, atomic, locked store. Foundation for everything; stdlib only so it tests without a venv.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `lib/seo-data/tokenstore.py`
|
||||||
|
- Create: `lib/seo-data/seo-data.test.sh`
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: nothing.
|
||||||
|
- Produces: `load`, `list_accounts`, `get_refresh_token`, `save_account` (signatures in File Structure) + CLI `list`/`set`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test** — create `lib/seo-data/seo-data.test.sh`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
# Deterministic tests for the seo-data engine (no network, no venv).
|
||||||
|
set -u
|
||||||
|
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||||
|
SD="$REPO/lib/seo-data"
|
||||||
|
PASS=0; FAIL=0
|
||||||
|
ok() { echo " PASS $1"; PASS=$((PASS+1)); }
|
||||||
|
no() { echo " FAIL $1 — $2"; FAIL=$((FAIL+1)); }
|
||||||
|
# assert stdout of a command contains / omits a fixed string
|
||||||
|
has() { if printf '%s' "$2" | grep -qF -- "$3"; then ok "$1"; else no "$1" "missing: $3"; fi; }
|
||||||
|
hasnt(){ if printf '%s' "$2" | grep -qF -- "$3"; then no "$1" "forbidden: $3"; else ok "$1"; fi; }
|
||||||
|
|
||||||
|
echo "── tokenstore ──"
|
||||||
|
TMP="$(mktemp -d)"; STORE="$TMP/tokens.json"
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$STORE" --label client-a \
|
||||||
|
--refresh-token RT_AAA --scopes https://www.googleapis.com/auth/webmasters.readonly \
|
||||||
|
--properties sc-domain:a.com,https://www.a.com/ >/dev/null
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$STORE" --label client-b \
|
||||||
|
--refresh-token RT_BBB --scopes https://www.googleapis.com/auth/webmasters.readonly \
|
||||||
|
--properties sc-domain:b.com >/dev/null
|
||||||
|
LIST="$(python3 "$SD/tokenstore.py" list --file "$STORE")"
|
||||||
|
has "list shows client-a" "$LIST" '"client-a"'
|
||||||
|
has "list shows client-b" "$LIST" '"client-b"'
|
||||||
|
has "list shows a property" "$LIST" 'sc-domain:a.com'
|
||||||
|
hasnt "list redacts refresh tokens" "$LIST" 'RT_AAA'
|
||||||
|
PERM="$(stat -c '%a' "$STORE")"
|
||||||
|
[ "$PERM" = "600" ] && ok "store file is 0600" || no "store file 0600" "got $PERM"
|
||||||
|
rm -rf "$TMP"
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "seo-data engine: $PASS pass, $FAIL fail"
|
||||||
|
[ "$FAIL" -eq 0 ]
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — `tokenstore.py` does not exist (`python3: can't open file`).
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement `lib/seo-data/tokenstore.py`** (stdlib only):
|
||||||
|
|
||||||
|
```python
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Label-keyed OAuth refresh-token store. Atomic writes under an fcntl lock.
|
||||||
|
No third-party deps — must run without the venv (used by the offline test path)."""
|
||||||
|
import argparse, fcntl, json, os, sys, tempfile
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
def load(path):
|
||||||
|
if not os.path.exists(path):
|
||||||
|
return {"version": 1, "accounts": {}}
|
||||||
|
with open(path, "r", encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
|
||||||
|
def list_accounts(path):
|
||||||
|
data = load(path)
|
||||||
|
return [
|
||||||
|
{"label": lbl, "properties": a.get("properties", []),
|
||||||
|
"granted_at": a.get("granted_at")}
|
||||||
|
for lbl, a in data.get("accounts", {}).items()
|
||||||
|
] # refresh_token intentionally omitted (redaction)
|
||||||
|
|
||||||
|
def get_refresh_token(path, label):
|
||||||
|
return load(path).get("accounts", {}).get(label, {}).get("refresh_token")
|
||||||
|
|
||||||
|
def save_account(path, label, refresh_token, scopes, properties):
|
||||||
|
os.makedirs(os.path.dirname(path), mode=0o700, exist_ok=True)
|
||||||
|
lock_path = path + ".lock"
|
||||||
|
with open(lock_path, "w") as lock:
|
||||||
|
fcntl.flock(lock, fcntl.LOCK_EX) # serialize concurrent connects
|
||||||
|
data = load(path)
|
||||||
|
data.setdefault("version", 1)
|
||||||
|
data.setdefault("accounts", {})
|
||||||
|
data["accounts"][label] = {
|
||||||
|
"refresh_token": refresh_token,
|
||||||
|
"scopes": scopes,
|
||||||
|
"granted_at": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"properties": properties,
|
||||||
|
}
|
||||||
|
fd, tmp = tempfile.mkstemp(dir=os.path.dirname(path), suffix=".tmp")
|
||||||
|
try:
|
||||||
|
with os.fdopen(fd, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(data, f, indent=2)
|
||||||
|
f.flush(); os.fsync(f.fileno())
|
||||||
|
os.chmod(tmp, 0o600)
|
||||||
|
os.replace(tmp, path) # atomic
|
||||||
|
finally:
|
||||||
|
if os.path.exists(tmp):
|
||||||
|
os.unlink(tmp)
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
sub = p.add_subparsers(dest="cmd", required=True)
|
||||||
|
pl = sub.add_parser("list"); pl.add_argument("--file", required=True)
|
||||||
|
ps = sub.add_parser("set")
|
||||||
|
for flag in ("--file", "--label", "--refresh-token"):
|
||||||
|
ps.add_argument(flag, required=True)
|
||||||
|
ps.add_argument("--scopes", default="")
|
||||||
|
ps.add_argument("--properties", default="")
|
||||||
|
args = p.parse_args()
|
||||||
|
if args.cmd == "list":
|
||||||
|
print(json.dumps({"status": "ok", "accounts": list_accounts(args.file)}))
|
||||||
|
else:
|
||||||
|
save_account(args.file, args.label, getattr(args, "refresh_token"),
|
||||||
|
[s for s in args.scopes.split(",") if s],
|
||||||
|
[x for x in args.properties.split(",") if x])
|
||||||
|
print(json.dumps({"status": "ok"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (5 tokenstore checks pass).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/tokenstore.py lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo-data): label-keyed atomic OAuth token store"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 2: CrUX fetch (`google_seo.py` — CrUX path)
|
||||||
|
|
||||||
|
Simplest data path (API key, no OAuth). Establishes the mock-mode + degrade + normalization pattern.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `lib/seo-data/google_seo.py`
|
||||||
|
- Create: `lib/seo-data/fixtures/crux_mobile.json`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append CrUX section)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: env `CRUX_API_KEY`, env `SEO_DATA_MOCK_DIR`.
|
||||||
|
- Produces: `crux(url, strategy='mobile') -> dict`; CLI `python3 google_seo.py crux --url … [--strategy …]`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the fixture** — `lib/seo-data/fixtures/crux_mobile.json` (shape of the CrUX API `record.metrics`):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"record":{"key":{"formFactor":"PHONE"},"metrics":{
|
||||||
|
"largest_contentful_paint":{"percentiles":{"p75":2100}},
|
||||||
|
"interaction_to_next_paint":{"percentiles":{"p75":180}},
|
||||||
|
"cumulative_layout_shift":{"percentiles":{"p75":"0.08"}}}}}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Write the failing test** — append to `lib/seo-data/seo-data.test.sh` before the final summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── crux (mock) ──"
|
||||||
|
CRUX_OK="$(SEO_DATA_MOCK_DIR="$REPO/lib/seo-data/fixtures" \
|
||||||
|
python3 "$SD/google_seo.py" crux --url https://ex.com --strategy mobile)"
|
||||||
|
has "crux status ok" "$CRUX_OK" '"status": "ok"'
|
||||||
|
has "crux lcp p75 mapped" "$CRUX_OK" '"lcp_p75_ms": 2100'
|
||||||
|
has "crux inp p75 mapped" "$CRUX_OK" '"inp_p75_ms": 180'
|
||||||
|
has "crux cls p75 mapped" "$CRUX_OK" '"cls_p75": 0.08'
|
||||||
|
CRUX_DEG="$(env -u CRUX_API_KEY -u SEO_DATA_MOCK_DIR \
|
||||||
|
python3 "$SD/google_seo.py" crux --url https://ex.com)"
|
||||||
|
has "crux degrades w/o key" "$CRUX_DEG" '"status": "degraded"'
|
||||||
|
has "crux degrade reason" "$CRUX_DEG" 'no_crux_key'
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — `google_seo.py` missing.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Implement the CrUX path** — create `lib/seo-data/google_seo.py` (lazy `requests` import; mock reads the fixture and runs the REAL normalizer):
|
||||||
|
|
||||||
|
```python
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""CrUX + GSC fetch → normalized JSON. Third-party imports are LAZY so mock and
|
||||||
|
degraded paths run stdlib-only (no venv, no network)."""
|
||||||
|
import argparse, json, os, sys
|
||||||
|
|
||||||
|
def _mock(name):
|
||||||
|
d = os.environ.get("SEO_DATA_MOCK_DIR")
|
||||||
|
if not d:
|
||||||
|
return None
|
||||||
|
path = os.path.join(d, name)
|
||||||
|
if not os.path.exists(path):
|
||||||
|
return None
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
|
||||||
|
def _norm_crux(raw):
|
||||||
|
m = raw["record"]["metrics"]
|
||||||
|
def p75(metric):
|
||||||
|
return m.get(metric, {}).get("percentiles", {}).get("p75")
|
||||||
|
out = {"status": "ok", "source": "crux"}
|
||||||
|
lcp = p75("largest_contentful_paint")
|
||||||
|
inp = p75("interaction_to_next_paint")
|
||||||
|
cls = p75("cumulative_layout_shift")
|
||||||
|
# Low-traffic origins often miss a metric (INP notably) — omit, don't crash.
|
||||||
|
if lcp is not None:
|
||||||
|
out["lcp_p75_ms"] = int(lcp)
|
||||||
|
if inp is not None:
|
||||||
|
out["inp_p75_ms"] = int(inp)
|
||||||
|
if cls is not None:
|
||||||
|
out["cls_p75"] = float(cls)
|
||||||
|
if len(out) == 2: # no metric at all
|
||||||
|
return {"status": "degraded", "reason": "no_field_data"}
|
||||||
|
return out
|
||||||
|
|
||||||
|
def _crux_query(key, body):
|
||||||
|
import requests # lazy
|
||||||
|
return requests.post(
|
||||||
|
"https://chromeuxreport.googleapis.com/v1/records:queryRecord?key=" + key,
|
||||||
|
json=body, timeout=20)
|
||||||
|
|
||||||
|
def crux(url, strategy="mobile"):
|
||||||
|
raw = _mock("crux_%s.json" % strategy)
|
||||||
|
if raw is None:
|
||||||
|
key = os.environ.get("CRUX_API_KEY")
|
||||||
|
if not key:
|
||||||
|
return {"status": "degraded", "reason": "no_crux_key"}
|
||||||
|
ff = "PHONE" if strategy == "mobile" else "DESKTOP"
|
||||||
|
r = _crux_query(key, {"url": url, "formFactor": ff})
|
||||||
|
if r.status_code == 404: # no page-level data → try origin-level
|
||||||
|
r = _crux_query(key, {"origin": url.rstrip("/"), "formFactor": ff})
|
||||||
|
if r.status_code == 404:
|
||||||
|
return {"status": "degraded", "reason": "no_field_data"}
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
return _norm_crux(raw)
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
sub = p.add_subparsers(dest="cmd", required=True)
|
||||||
|
pc = sub.add_parser("crux")
|
||||||
|
pc.add_argument("--url", required=True)
|
||||||
|
pc.add_argument("--strategy", default="mobile", choices=["mobile", "desktop"])
|
||||||
|
pc.add_argument("--store", default=None) # accepted+ignored: uniform fetch.sh dispatch
|
||||||
|
args = p.parse_args()
|
||||||
|
try:
|
||||||
|
if args.cmd == "crux":
|
||||||
|
print(json.dumps(crux(args.url, args.strategy), indent=2))
|
||||||
|
except Exception:
|
||||||
|
# Fail-open data contract: ANY unexpected error (HTTP 403/5xx, DNS,
|
||||||
|
# timeout) degrades with exit 0 — never a traceback, never empty stdout.
|
||||||
|
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (tokenstore + 6 CrUX checks).
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/google_seo.py lib/seo-data/fixtures/crux_mobile.json lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo-data): CrUX field-data fetch with mock mode and graceful degrade"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 3: GSC fetch (`google_seo.py` — queries + inspect)
|
||||||
|
|
||||||
|
Adds Search Analytics + URL Inspection with OAuth refresh (lazy) reusing `tokenstore`.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `lib/seo-data/google_seo.py` (add `queries`, `inspect`, `_gsc_session`, extend CLI)
|
||||||
|
- Create: `lib/seo-data/fixtures/gsc_queries.json`, `lib/seo-data/fixtures/gsc_inspect.json`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append GSC section)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `tokenstore.get_refresh_token`, env `GOOGLE_OAUTH_CLIENT_ID/SECRET`, `SEO_DATA_MOCK_DIR`.
|
||||||
|
- Produces: `queries(store_path, account, property, days=90, dim='query')`, `inspect(store_path, account, property, url)`; CLI `queries`/`inspect`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write fixtures**
|
||||||
|
|
||||||
|
`lib/seo-data/fixtures/gsc_queries.json` (Search Analytics `rows` shape):
|
||||||
|
```json
|
||||||
|
{"rows":[
|
||||||
|
{"keys":["plombier paris"],"clicks":40,"impressions":900,"ctr":0.044,"position":6.3},
|
||||||
|
{"keys":["urgence fuite"],"clicks":5,"impressions":1200,"ctr":0.004,"position":8.9}]}
|
||||||
|
```
|
||||||
|
`lib/seo-data/fixtures/gsc_inspect.json` (URL Inspection shape):
|
||||||
|
```json
|
||||||
|
{"inspectionResult":{"indexStatusResult":{
|
||||||
|
"verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Write the failing test** — append before the summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── gsc (mock) ──"
|
||||||
|
MOCK="$REPO/lib/seo-data/fixtures"
|
||||||
|
TMP2="$(mktemp -d)"; S2="$TMP2/tokens.json"
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$S2" --label client-a --refresh-token RT \
|
||||||
|
--scopes https://www.googleapis.com/auth/webmasters.readonly --properties sc-domain:ex.com >/dev/null
|
||||||
|
Q="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" queries \
|
||||||
|
--store "$S2" --account client-a --property sc-domain:ex.com --days 90)"
|
||||||
|
has "queries ok" "$Q" '"status": "ok"'
|
||||||
|
has "queries row key" "$Q" 'plombier paris'
|
||||||
|
has "queries position field" "$Q" '"position": 6.3'
|
||||||
|
I="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" inspect \
|
||||||
|
--store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)"
|
||||||
|
has "inspect indexed true" "$I" '"indexed": true'
|
||||||
|
DEG="$(env -u SEO_DATA_MOCK_DIR python3 "$SD/google_seo.py" queries \
|
||||||
|
--store "$TMP2/none.json" --account nobody --property sc-domain:ex.com)"
|
||||||
|
has "gsc degrades w/o creds" "$DEG" '"status": "degraded"'
|
||||||
|
has "gsc degrade reason" "$DEG" 'no_credentials'
|
||||||
|
rm -rf "$TMP2"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — `queries`/`inspect` not implemented (argparse error / AttributeError).
|
||||||
|
|
||||||
|
- [ ] **Step 4: Implement** — add to `google_seo.py`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _gsc_session(store_path, account):
|
||||||
|
"""Return an authorized requests.Session or a degrade dict. Lazy imports."""
|
||||||
|
rt = None
|
||||||
|
if store_path and account:
|
||||||
|
import tokenstore # local module, stdlib
|
||||||
|
rt = tokenstore.get_refresh_token(store_path, account)
|
||||||
|
cid = os.environ.get("GOOGLE_OAUTH_CLIENT_ID")
|
||||||
|
csec = os.environ.get("GOOGLE_OAUTH_CLIENT_SECRET")
|
||||||
|
if not (rt and cid and csec):
|
||||||
|
return {"status": "degraded", "reason": "no_credentials"}
|
||||||
|
from google.oauth2.credentials import Credentials # lazy
|
||||||
|
from google.auth.transport.requests import AuthorizedSession, Request
|
||||||
|
creds = Credentials(None, refresh_token=rt, client_id=cid, client_secret=csec,
|
||||||
|
token_uri="https://oauth2.googleapis.com/token",
|
||||||
|
scopes=["https://www.googleapis.com/auth/webmasters.readonly"])
|
||||||
|
try:
|
||||||
|
creds.refresh(Request())
|
||||||
|
except Exception as e:
|
||||||
|
# Only a real RefreshError means re-consent; a network blip must NOT
|
||||||
|
# send the user back through OAuth.
|
||||||
|
from google.auth.exceptions import RefreshError # lazy
|
||||||
|
reason = "token_revoked" if isinstance(e, RefreshError) else "network_error"
|
||||||
|
return {"status": "degraded", "reason": reason}
|
||||||
|
return AuthorizedSession(creds)
|
||||||
|
|
||||||
|
def _norm_queries(raw, dim):
|
||||||
|
return {"status": "ok", "source": "gsc", "dimension": dim, "rows": [
|
||||||
|
{"key": r["keys"][0], "clicks": r.get("clicks", 0),
|
||||||
|
"impressions": r.get("impressions", 0), "ctr": r.get("ctr", 0),
|
||||||
|
"position": r.get("position")}
|
||||||
|
for r in raw.get("rows", [])]}
|
||||||
|
|
||||||
|
def queries(store_path, account, property, days=90, dim="query"):
|
||||||
|
raw = _mock("gsc_queries.json")
|
||||||
|
if raw is None:
|
||||||
|
sess = _gsc_session(store_path, account)
|
||||||
|
if isinstance(sess, dict):
|
||||||
|
return sess
|
||||||
|
import datetime as _dt
|
||||||
|
end = _dt.date.today(); start = end - _dt.timedelta(days=days)
|
||||||
|
import urllib.parse
|
||||||
|
url = ("https://searchconsole.googleapis.com/webmasters/v3/sites/"
|
||||||
|
+ urllib.parse.quote(property, safe="") + "/searchAnalytics/query")
|
||||||
|
r = sess.post(url, json={"startDate": start.isoformat(), "endDate": end.isoformat(),
|
||||||
|
"dimensions": [dim], "rowLimit": 100}, timeout=30)
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
return _norm_queries(raw, dim)
|
||||||
|
|
||||||
|
def inspect(store_path, account, property, url):
|
||||||
|
raw = _mock("gsc_inspect.json")
|
||||||
|
if raw is None:
|
||||||
|
sess = _gsc_session(store_path, account)
|
||||||
|
if isinstance(sess, dict):
|
||||||
|
return sess
|
||||||
|
r = sess.post("https://searchconsole.googleapis.com/v1/urlInspection/index:inspect",
|
||||||
|
json={"inspectionUrl": url, "siteUrl": property}, timeout=30)
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
isr = raw["inspectionResult"]["indexStatusResult"]
|
||||||
|
return {"status": "ok", "source": "gsc",
|
||||||
|
"indexed": isr.get("verdict") == "PASS",
|
||||||
|
"coverage": isr.get("coverageState"),
|
||||||
|
"last_crawl": isr.get("lastCrawlTime")}
|
||||||
|
```
|
||||||
|
Extend `_cli()` (add subparsers `queries` and `inspect`, each with `--store --account --property`, plus `--days`/`--dim` for queries and `--url` for inspect; dispatch to the functions and `print(json.dumps(..., indent=2))` **inside the existing top-level `try/except`** from Task 2 — the fail-open contract covers every subcommand). Ensure the script's dir is importable for `import tokenstore` (add `sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))` at top).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (tokenstore + CrUX + 7 GSC checks).
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/google_seo.py lib/seo-data/fixtures/gsc_queries.json lib/seo-data/fixtures/gsc_inspect.json lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo-data): GSC Search Analytics + URL Inspection with lazy OAuth refresh"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 4: Bash entrypoint (`fetch.sh`)
|
||||||
|
|
||||||
|
The stable CLI the analyzers call. Sources env, picks python (venv else system), dispatches, guarantees degrade-exit-0 and redaction.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `lib/seo-data/fetch.sh`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append fetch.sh section)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `~/.claude/.env` (canonical), `google_seo.py`, `tokenstore.py`, optional `~/.claude/.venv-seo-data/`.
|
||||||
|
- Produces: `fetch.sh {accounts|crux|queries|inspect} [flags]` → JSON stdout, exit 0 on ok/degrade, exit 2 on bad usage.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test** — append before the summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── fetch.sh ──"
|
||||||
|
FETCH="$SD/fetch.sh"
|
||||||
|
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
|
||||||
|
# on a machine with a live CRUX_API_KEY the degrade tests would hit the network.
|
||||||
|
NOENV=/dev/null
|
||||||
|
ACC="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_STORE=/nonexistent/tokens.json bash "$FETCH" accounts)"
|
||||||
|
has "accounts empty is ok json" "$ACC" '"accounts": []'
|
||||||
|
CR="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_MOCK_DIR="$MOCK" bash "$FETCH" crux --url https://ex.com)"
|
||||||
|
has "fetch crux ok" "$CR" '"status": "ok"'
|
||||||
|
SEO_DATA_ENV_FILE=$NOENV bash "$FETCH" bogus-subcmd >/dev/null 2>&1; RC=$?
|
||||||
|
[ "$RC" = "2" ] && ok "bad subcmd exit 2" || no "bad subcmd exit 2" "got $RC"
|
||||||
|
DG="$(SEO_DATA_ENV_FILE=$NOENV env -u SEO_DATA_MOCK_DIR -u CRUX_API_KEY bash "$FETCH" crux --url https://ex.com)"; RC=$?
|
||||||
|
has "degrade json" "$DG" '"status": "degraded"'
|
||||||
|
[ "$RC" = "0" ] && ok "degrade exit 0" || no "degrade exit 0" "got $RC"
|
||||||
|
hasnt "no secret echoed" "$DG" 'RT_'
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — `fetch.sh` missing.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement `lib/seo-data/fetch.sh`:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
# Stable entrypoint for the seo-data engine. JSON on stdout; exit 0 on ok/degrade,
|
||||||
|
# exit 2 on bad usage. Never prints secrets.
|
||||||
|
set -uo pipefail
|
||||||
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
ENV_FILE="${SEO_DATA_ENV_FILE:-${HOME}/.claude/.env}" # canonical; tests override to /dev/null
|
||||||
|
STORE="${SEO_DATA_STORE:-${HOME}/.claude/seo-data/tokens.json}"
|
||||||
|
VENV_PY="${HOME}/.claude/.venv-seo-data/bin/python3"
|
||||||
|
|
||||||
|
# Library stderr must never leak a secret into agent context — suppress it
|
||||||
|
# globally unless explicitly debugging (SEO_DATA_DEBUG=1 restores it).
|
||||||
|
[ -n "${SEO_DATA_DEBUG:-}" ] || exec 2>/dev/null
|
||||||
|
|
||||||
|
# Load secrets quietly (sourced, never echoed).
|
||||||
|
if [ -f "$ENV_FILE" ]; then
|
||||||
|
set -a; # shellcheck source=/dev/null
|
||||||
|
. "$ENV_FILE"; set +a
|
||||||
|
fi
|
||||||
|
# Prefer the isolated venv (has google-auth); fall back to system python3 for
|
||||||
|
# stdlib-only paths (accounts / mock / degrade).
|
||||||
|
PY="python3"; [ -x "$VENV_PY" ] && PY="$VENV_PY"
|
||||||
|
|
||||||
|
cmd="${1:-}"; shift || true
|
||||||
|
case "$cmd" in
|
||||||
|
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
||||||
|
crux|queries|inspect)
|
||||||
|
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
||||||
|
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect} [flags]"}'
|
||||||
|
exit 2 ;;
|
||||||
|
esac
|
||||||
|
```
|
||||||
|
Notes: `crux` accepts and ignores `--store` (already wired in Task 2's `_cli`) so `fetch.sh` dispatches uniformly. The global `exec 2>/dev/null` (unless `SEO_DATA_DEBUG=1`) plus `google_seo.py`'s top-level `try/except` together guarantee: JSON always on stdout, never a traceback, never a leaked secret, exit 0 on every data-path outcome.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (all prior + 6 fetch.sh checks).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/fetch.sh lib/seo-data/google_seo.py lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo-data): fetch.sh entrypoint with venv/system fallback and redaction"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 5: OAuth consent (`connect.py`) + pinned deps
|
||||||
|
|
||||||
|
One-time interactive consent + `sites.list` discovery + persist. Browser flow is manually verified; the persist + label logic is unit-tested.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `lib/seo-data/connect.py`
|
||||||
|
- Create: `lib/seo-data/requirements.txt`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append persist test)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: env `GOOGLE_OAUTH_CLIENT_ID/SECRET`, `tokenstore.save_account`.
|
||||||
|
- Produces: `run_consent`, `discover_properties`, `persist`; CLI `python3 connect.py --label LABEL`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write `requirements.txt`** (pinned; versions current as of 2026-07 — the implementer verifies latest patch at execution):
|
||||||
|
|
||||||
|
```
|
||||||
|
google-auth==2.40.0
|
||||||
|
google-auth-oauthlib==1.2.2
|
||||||
|
requests==2.32.4
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Write the failing test** — append before the summary (tests only the offline-safe `persist`, via the tokenstore it wraps):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── connect (persist, offline) ──"
|
||||||
|
TMP3="$(mktemp -d)"; S3="$TMP3/tokens.json"
|
||||||
|
python3 -c "import sys; sys.path.insert(0,'$SD'); import connect; \
|
||||||
|
connect.persist('$S3','client-x','RT_X',['https://www.googleapis.com/auth/webmasters.readonly'],['sc-domain:x.com'])"
|
||||||
|
L3="$(python3 "$SD/tokenstore.py" list --file "$S3")"
|
||||||
|
has "connect.persist wrote label" "$L3" '"client-x"'
|
||||||
|
has "connect.persist wrote prop" "$L3" 'sc-domain:x.com'
|
||||||
|
hasnt "connect.persist redacts" "$L3" 'RT_X'
|
||||||
|
rm -rf "$TMP3"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — `connect` module / `persist` missing.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Implement `lib/seo-data/connect.py`:**
|
||||||
|
|
||||||
|
```python
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""One-time OAuth consent + GSC property discovery + persist. Third-party imports
|
||||||
|
are lazy so `persist` is testable stdlib-only."""
|
||||||
|
import argparse, os, sys
|
||||||
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
import tokenstore
|
||||||
|
|
||||||
|
SCOPES = ["https://www.googleapis.com/auth/webmasters.readonly"]
|
||||||
|
|
||||||
|
def run_consent(client_id, client_secret):
|
||||||
|
from google_auth_oauthlib.flow import InstalledAppFlow # lazy
|
||||||
|
cfg = {"installed": {"client_id": client_id, "client_secret": client_secret,
|
||||||
|
"auth_uri": "https://accounts.google.com/o/oauth2/auth",
|
||||||
|
"token_uri": "https://oauth2.googleapis.com/token",
|
||||||
|
"redirect_uris": ["http://localhost"]}}
|
||||||
|
flow = InstalledAppFlow.from_client_config(cfg, scopes=SCOPES)
|
||||||
|
creds = flow.run_local_server(port=0) # opens browser, one-time consent
|
||||||
|
if not creds.refresh_token:
|
||||||
|
raise SystemExit("No refresh token returned. Revoke prior grant and retry.")
|
||||||
|
return creds.refresh_token
|
||||||
|
|
||||||
|
def discover_properties(refresh_token, client_id, client_secret):
|
||||||
|
from google.oauth2.credentials import Credentials
|
||||||
|
from google.auth.transport.requests import AuthorizedSession, Request
|
||||||
|
creds = Credentials(None, refresh_token=refresh_token, client_id=client_id,
|
||||||
|
client_secret=client_secret,
|
||||||
|
token_uri="https://oauth2.googleapis.com/token", scopes=SCOPES)
|
||||||
|
creds.refresh(Request())
|
||||||
|
r = AuthorizedSession(creds).get(
|
||||||
|
"https://searchconsole.googleapis.com/webmasters/v3/sites", timeout=30)
|
||||||
|
r.raise_for_status()
|
||||||
|
return [e["siteUrl"] for e in r.json().get("siteEntry", [])]
|
||||||
|
|
||||||
|
def persist(store_path, label, refresh_token, scopes, properties):
|
||||||
|
tokenstore.save_account(store_path, label, refresh_token, scopes, properties)
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
p.add_argument("--label", required=True)
|
||||||
|
p.add_argument("--store", default=os.path.expanduser("~/.claude/seo-data/tokens.json"))
|
||||||
|
args = p.parse_args()
|
||||||
|
cid = os.environ.get("GOOGLE_OAUTH_CLIENT_ID")
|
||||||
|
csec = os.environ.get("GOOGLE_OAUTH_CLIENT_SECRET")
|
||||||
|
if not (cid and csec):
|
||||||
|
raise SystemExit("Set GOOGLE_OAUTH_CLIENT_ID/SECRET in ~/.claude/.env first.")
|
||||||
|
existing = {a["label"] for a in tokenstore.list_accounts(args.store)}
|
||||||
|
if args.label in existing:
|
||||||
|
ans = input("Label '%s' exists. Overwrite? [y/N] " % args.label).strip().lower()
|
||||||
|
if ans != "y":
|
||||||
|
raise SystemExit("Aborted.")
|
||||||
|
rt = run_consent(cid, csec)
|
||||||
|
props = discover_properties(rt, cid, csec)
|
||||||
|
persist(args.store, args.label, rt, SCOPES, props)
|
||||||
|
print("Connected '%s'. Properties: %s" % (args.label, ", ".join(props) or "(none)"))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (all prior + 3 persist checks).
|
||||||
|
|
||||||
|
- [ ] **Step 6: Manual verification (documented, not automated)** — after Task 6 wires `make seo-connect`: run it once against a real GCP OAuth client, confirm the browser consent completes, the store gains the label with discovered properties, and a second `fetch.sh queries` runs non-interactively.
|
||||||
|
|
||||||
|
- [ ] **Step 7: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/connect.py lib/seo-data/requirements.txt lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo-data): OAuth consent + property discovery + pinned deps"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 6: Install / deploy wiring
|
||||||
|
|
||||||
|
`.env.example`, `Makefile seo-connect`, `install.sh` step, `doctor.sh` check, gitleaks allowlist, gitignore. All content-locked by the bash test.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `.env.example`, `Makefile`, `install.sh`, `doctor.sh`, `.gitleaks.toml`, `.gitignore`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append wiring locks)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `lib/seo-data/{connect.py,requirements.txt}`.
|
||||||
|
- Produces: `make seo-connect`; doctor check; allowlisted store path.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test** — append before the summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── wiring locks ──"
|
||||||
|
tf() { if grep -qF -- "$3" "$2" 2>/dev/null; then ok "$1"; else no "$1" "missing: $3"; fi; }
|
||||||
|
tf "env.example client id" "$REPO/.env.example" "GOOGLE_OAUTH_CLIENT_ID="
|
||||||
|
tf "env.example crux key" "$REPO/.env.example" "CRUX_API_KEY="
|
||||||
|
tf "makefile seo-connect" "$REPO/Makefile" "seo-connect:"
|
||||||
|
tf "makefile discovers test" "$REPO/Makefile" "lib/seo-data/*.test.sh"
|
||||||
|
tf "install prompts connect" "$REPO/install.sh" "make seo-connect"
|
||||||
|
tf "doctor checks seo-data" "$REPO/doctor.sh" "seo-data"
|
||||||
|
tf "gitleaks allowlist store" "$REPO/.gitleaks.toml" "seo-data/tokens.json"
|
||||||
|
tf "gitignore venv" "$REPO/.gitignore" ".venv-seo-data"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — none of the 8 locks present yet.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Apply the wiring edits**
|
||||||
|
|
||||||
|
`.env.example` — append:
|
||||||
|
```
|
||||||
|
# ── Google SEO data layer (lib/seo-data) — used by /seo FULL ──
|
||||||
|
# OAuth Desktop client: GCP console → APIs & Services → Credentials → OAuth client (Desktop).
|
||||||
|
# Scope requested at consent: webmasters.readonly. One-time setup: make seo-connect
|
||||||
|
GOOGLE_OAUTH_CLIENT_ID=<your-client-id.apps.googleusercontent.com>
|
||||||
|
GOOGLE_OAUTH_CLIENT_SECRET=<your-client-secret>
|
||||||
|
# CrUX + PageSpeed API key (GCP console → Credentials → API key, restricted to those APIs).
|
||||||
|
# Get it: https://developer.chrome.com/docs/crux/api
|
||||||
|
CRUX_API_KEY=<your-crux-api-key>
|
||||||
|
```
|
||||||
|
|
||||||
|
`Makefile` — add target + `.PHONY`:
|
||||||
|
```make
|
||||||
|
seo-connect: ## Connect a Google account for /seo FULL (creates venv, OAuth consent)
|
||||||
|
@python3 -m venv "$$HOME/.claude/.venv-seo-data"
|
||||||
|
@"$$HOME/.claude/.venv-seo-data/bin/pip" install -q -r lib/seo-data/requirements.txt
|
||||||
|
@bash -c 'read -r -p "Label for this account (e.g. client-a): " label; \
|
||||||
|
"$$HOME/.claude/.venv-seo-data/bin/python3" lib/seo-data/connect.py --label "$$label"'
|
||||||
|
```
|
||||||
|
(Add `seo-connect` to the `.PHONY:` line at the top of the Makefile.)
|
||||||
|
|
||||||
|
Also extend the existing `test:` target so `make test` discovers the engine test. Change its loop line:
|
||||||
|
```make
|
||||||
|
@fail=0; for t in lib/tests/*.test.sh lib/gitflow-test.sh lib/tests/run-*.sh; do \
|
||||||
|
```
|
||||||
|
to add `lib/seo-data/*.test.sh`:
|
||||||
|
```make
|
||||||
|
@fail=0; for t in lib/tests/*.test.sh lib/seo-data/*.test.sh lib/gitflow-test.sh lib/tests/run-*.sh; do \
|
||||||
|
```
|
||||||
|
(Editing the Makefile is allowed — it is not a config-protection-gated file. This *adds* test discovery; it does not weaken any gate.)
|
||||||
|
|
||||||
|
`install.sh` — after the `bash "$REPO/link.sh"` block (§5, ~line 107), before plugins:
|
||||||
|
```bash
|
||||||
|
# ── 5b. Optional: connect a Google account for /seo FULL ──
|
||||||
|
echo ""
|
||||||
|
if [ -f "$HOME/.claude/seo-data/tokens.json" ]; then
|
||||||
|
ok "seo-data: a Google account is already connected"
|
||||||
|
else
|
||||||
|
info "SEO data layer (GSC + CrUX) is optional. To enable real Search Console"
|
||||||
|
info "data in /seo FULL, add GOOGLE_OAUTH_* + CRUX_API_KEY to ~/.claude/.env,"
|
||||||
|
info "then run: make seo-connect"
|
||||||
|
fi
|
||||||
|
```
|
||||||
|
|
||||||
|
`doctor.sh` — add a check block (non-fatal, canonical env, mirrors existing WARN style):
|
||||||
|
```bash
|
||||||
|
echo "── seo-data (GSC/CrUX) ──"
|
||||||
|
if grep -qE '^[[:space:]]*(export[[:space:]]+)?CRUX_API_KEY=.' "$HOME/.claude/.env" 2>/dev/null; then
|
||||||
|
ok "CRUX_API_KEY present"
|
||||||
|
else
|
||||||
|
warn "CRUX_API_KEY absent in ~/.claude/.env — /seo FULL falls back to lab PageSpeed"
|
||||||
|
fi
|
||||||
|
if [ -f "$HOME/.claude/seo-data/tokens.json" ]; then
|
||||||
|
ok "seo-data: $(python3 "$REPO/lib/seo-data/tokenstore.py" list --file "$HOME/.claude/seo-data/tokens.json" | grep -o '"label"' | wc -l) account(s) connected"
|
||||||
|
else
|
||||||
|
warn "seo-data: no Google account connected (run: make seo-connect) — GSC data disabled"
|
||||||
|
fi
|
||||||
|
```
|
||||||
|
(Use the same `ok`/`warn` helpers doctor.sh already defines; if they differ, match its local names.)
|
||||||
|
|
||||||
|
`.gitleaks.toml` — add to `[allowlist].paths` (after the `.env` entry):
|
||||||
|
```toml
|
||||||
|
# seo-data OAuth token store — legitimate local secret (like ~/.claude/.env),
|
||||||
|
# 0600, outside git. Allowlisted so `make scan-secrets` doesn't flag the vault.
|
||||||
|
'''(^|/)\.claude/seo-data/tokens\.json$''',
|
||||||
|
```
|
||||||
|
|
||||||
|
`.gitignore` — after the `.env*` block (~line 114):
|
||||||
|
```
|
||||||
|
# seo-data engine local artifacts (live under ~/.claude, never committed)
|
||||||
|
.venv-seo-data/
|
||||||
|
seo-data/tokens.json
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (all prior + 8 wiring locks). Then `make test` — the whole suite still green (now including the seo-data engine test via the new glob).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add .env.example Makefile install.sh doctor.sh .gitleaks.toml .gitignore lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "chore(seo-data): install/make/doctor wiring + gitleaks allowlist for token store"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 7: Analyzer + skill integration
|
||||||
|
|
||||||
|
Make `/seo` FULL actually consume the engine: account selection in STEP 0, CrUX field data + a GSC performance subsection in the analyzer, automation-catalog entry.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `skills/seo/SKILL.md` (STEP 0 — account/property selection, FULL only)
|
||||||
|
- Modify: `agents/seo-analyzer.md` (STEP 4 CWV terrain via `fetch.sh crux`; new "Performance GSC" subsection via `fetch.sh queries`/`inspect`; STEP 9 Technical axis note)
|
||||||
|
- Modify: `agents/resources/automation-catalog.md` (GSC OAuth connection entry)
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append integration locks)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `lib/seo-data/fetch.sh` CLI contract.
|
||||||
|
- Produces: analyzer output enriched with real GSC/CrUX; passes `(account, property)` explicitly.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test** — append before the summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── integration locks ──"
|
||||||
|
tf "skill step0 account select" "$REPO/skills/seo/SKILL.md" "COMPTE GOOGLE"
|
||||||
|
tf "analyzer calls fetch crux" "$REPO/agents/seo-analyzer.md" "fetch.sh crux"
|
||||||
|
tf "analyzer calls fetch queries" "$REPO/agents/seo-analyzer.md" "fetch.sh queries"
|
||||||
|
tf "analyzer gsc subsection" "$REPO/agents/seo-analyzer.md" "Performance GSC"
|
||||||
|
tf "catalog gsc oauth entry" "$REPO/agents/resources/automation-catalog.md" "make seo-connect"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run it, verify red**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: FAIL — 5 integration locks absent.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Apply the integration edits** (concrete anchors from the /analyze report):
|
||||||
|
|
||||||
|
`skills/seo/SKILL.md` — in **STEP 0**, after the "Audit depth" block, add a FULL-only account-selection block (main loop, interactive) exactly as specified in spec §6, opening with the line `COMPTE GOOGLE pour cet audit FULL :`, listing connected accounts from `bash ~/.claude/lib/seo-data/fetch.sh accounts` (**tilde path mandatory** — skills run from the audited project's directory, not the claude-config repo), an option to run `make seo-connect`, and an "Ignore" option. Record the chosen `(account, property)` in the shared context block and pass it into **both** analyzer dispatch prompts (STEP 1) under `BUSINESS CONTEXT` as `GSC account: <label>` / `GSC property: <property>`.
|
||||||
|
|
||||||
|
`agents/seo-analyzer.md` — in **STEP 4 → Core Web Vitals** (currently the anonymous PageSpeed curl, ~lines 238-259), add before the PageSpeed call:
|
||||||
|
```
|
||||||
|
When a GSC account+property were passed in context, fetch CrUX field data first:
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh crux --url "https://$DOMAIN" --strategy mobile
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh crux --url "https://$DOMAIN" --strategy desktop
|
||||||
|
If status=ok, use lcp_p75_ms / inp_p75_ms / cls_p75 as the PRIMARY CWV figures
|
||||||
|
(75th pct, real users). Keep the PageSpeed lab run as a SECONDARY diagnostic.
|
||||||
|
If status=degraded, fall back to PageSpeed lab only (current behavior).
|
||||||
|
```
|
||||||
|
Then add a new subsection **"Performance GSC (90 j)"** in STEP 4 (FULL only, when account+property present):
|
||||||
|
```
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh queries --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --days 90 --dim query
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh inspect --account "$GSC_ACCOUNT" --property "$GSC_PROPERTY" --url "https://$DOMAIN/"
|
||||||
|
Report: top queries; flag QUICK WINS = rows with position between 4 and 10 AND
|
||||||
|
high impressions (candidates to push onto page 1). Report index coverage from inspect.
|
||||||
|
All emitted into SEO.md §2 and §8. If status=degraded → note it and emit the §11
|
||||||
|
user action "Connecter GSC: make seo-connect".
|
||||||
|
```
|
||||||
|
In **STEP 9** scoring, add a note on the Technical axis: "CWV scored on CrUX field data when available; else lab PageSpeed." In **STEP 3 (plugin check)**, add the graceful-degradation line for GSC/CrUX creds (READY | DEGRADED).
|
||||||
|
|
||||||
|
`agents/resources/automation-catalog.md` — under the existing Google Search Console entry (~line 69), add:
|
||||||
|
```
|
||||||
|
- **Connexion GSC pour /seo (données réelles)** — `make seo-connect` (depuis le repo
|
||||||
|
claude-config, une fois par compte) : consentement OAuth lecture seule
|
||||||
|
(webmasters.readonly), stocke un refresh token local (0600). Ensuite /seo FULL lit
|
||||||
|
requêtes/positions/indexation sans réinvite.
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run it, verify green**
|
||||||
|
|
||||||
|
Run: `bash lib/seo-data/seo-data.test.sh`
|
||||||
|
Expected: PASS (all prior + 5 integration locks). Run `make test` — full suite green.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add skills/seo/SKILL.md agents/seo-analyzer.md agents/resources/automation-catalog.md lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "feat(seo): wire GSC+CrUX data into /seo FULL (STEP 0 account select, CWV field, GSC perf)"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 8: Engine README
|
||||||
|
|
||||||
|
Document the engine's contract so future maintainers (and the analyzer) have a single reference.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `lib/seo-data/README.md`
|
||||||
|
- Modify: `lib/seo-data/seo-data.test.sh` (append doc lock)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: nothing.
|
||||||
|
- Produces: `lib/seo-data/README.md`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test** — append before the summary:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
echo "── readme lock ──"
|
||||||
|
tf "readme documents fetch.sh" "$REPO/lib/seo-data/README.md" "fetch.sh"
|
||||||
|
tf "readme documents seo-connect" "$REPO/lib/seo-data/README.md" "make seo-connect"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run it, verify red** — Run: `bash lib/seo-data/seo-data.test.sh` — Expected: FAIL (README missing).
|
||||||
|
|
||||||
|
- [ ] **Step 3: Write `lib/seo-data/README.md`** — cover: purpose (real GSC+CrUX for /seo FULL), setup (`make seo-connect` + the 3 `.env` keys), the `fetch.sh` subcommand contract (copy §9 of the spec), the token store location + security notes (0600, gitleaks allowlist, scope webmasters.readonly), graceful-degradation behavior, and how to run the test (`make test`). Concrete, no placeholders.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run it, verify green** — Run: `bash lib/seo-data/seo-data.test.sh` — Expected: PASS (all locks).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add lib/seo-data/README.md lib/seo-data/seo-data.test.sh
|
||||||
|
git commit -m "docs(seo-data): engine usage + security contract README"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Self-Review
|
||||||
|
|
||||||
|
**1. Spec coverage** — every spec section maps to a task:
|
||||||
|
- §4 architecture → Tasks 1-5 (engine files) + 7 (consumers).
|
||||||
|
- §5 auth/secrets (OAuth, .env vault, keyed store, gitleaks) → Task 1 (store), 5 (OAuth), 6 (.env.example + gitleaks + gitignore).
|
||||||
|
- §6 account selection STEP 0 → Task 7.
|
||||||
|
- §7 concurrency (explicit args, read-only audits, atomic locked writes) → Task 1 (atomic/lock), 4 (explicit args), 7 (pass-through).
|
||||||
|
- §8 data mapping (CrUX field, GSC queries/inspect, position 4-10) → Tasks 2, 3, 7.
|
||||||
|
- §9 fetch.sh contract → Task 4.
|
||||||
|
- §10 degradation/redaction → Tasks 2, 3, 4 (degrade+exit0+redaction tests).
|
||||||
|
- §11 install (env.example, install.sh, Makefile, doctor, gitleaks, gitignore) → Task 6.
|
||||||
|
- §12 tests → woven through every task (bash, fixtures, no network).
|
||||||
|
- §14 file touch-list → matches Tasks 1-8 exactly.
|
||||||
|
- §13 YAGNI (no GA4/Ads/Indexing, geo-analyzer untouched, one property/audit) → respected (no such tasks).
|
||||||
|
|
||||||
|
**2. Placeholder scan** — no "TBD/TODO". Tasks 3, 7, 8 use precise prose for markdown edits with exact anchor strings + verbatim content to insert (not "add error handling"). All code steps show code. `requirements.txt` versions flagged "verify latest patch at execution" (a real instruction, not a placeholder).
|
||||||
|
|
||||||
|
**3. Type consistency** — signatures identical across tasks: `save_account(path,label,refresh_token,scopes,properties)`, `get_refresh_token(path,label)`, `list_accounts(path)`, `crux(url,strategy)`, `queries(store_path,account,property,days,dim)`, `inspect(store_path,account,property,url)`, `persist(...)` delegates to `save_account`. `fetch.sh` passes `--store` uniformly (crux ignores it — noted in Task 4). `SEO_DATA_MOCK_DIR` / `SEO_DATA_STORE` env names consistent.
|
||||||
|
|
||||||
|
**Known execution caveats (not gaps):** the OAuth browser consent (Task 5) and live API calls can't be unit-tested — covered by the Task 5 manual-verification step. The `.test.sh` accumulates sections across tasks; each task appends before the final summary block.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Execution Handoff
|
||||||
|
|
||||||
|
Plan complete and saved to `docs/superpowers/plans/2026-07-09-gsc-crux-data-layer.md`. Two execution options:
|
||||||
|
|
||||||
|
**1. Subagent-Driven (recommended)** — a fresh subagent per task, review between tasks, fast iteration.
|
||||||
|
**2. Inline Execution** — execute tasks in this session with checkpoints for review.
|
||||||
|
|
||||||
|
Which approach?
|
||||||
@@ -0,0 +1,383 @@
|
|||||||
|
# Spec — Couche data GSC + CrUX pour `/seo` (+`/geo`) FULL
|
||||||
|
|
||||||
|
- **Date** : 2026-07-09
|
||||||
|
- **Statut** : Design validé — prêt pour `/writing-plans`
|
||||||
|
- **Auteur** : Bastien Chanot (design assisté)
|
||||||
|
- **Repo** : claude-config (`~/Documents/claude`, symlinké dans `~/.claude` via `link.sh`)
|
||||||
|
- **Cycle de vie** : document de travail **transitoire** — à supprimer une fois la feature livrée,
|
||||||
|
documentée (`/document-release` ou `/doc`) et capitalisée (`decisions.md`). Ne pas conserver à long terme.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Contexte & objectif
|
||||||
|
|
||||||
|
L'audit comparatif entre les skills perso `/seo` + `/geo` et l'outil marketplace
|
||||||
|
`agricidaniel/claude-seo` a isolé **un seul gap structurel** : les skills perso ne
|
||||||
|
peuvent pas lire la **donnée Google réelle** d'un site (requêtes/positions/impressions
|
||||||
|
de la Search Console, statut d'indexation, Core Web Vitals **terrain**). Ils se limitent
|
||||||
|
à l'API PageSpeed anonyme (données *labo*) et à `WebSearch`.
|
||||||
|
|
||||||
|
**Objectif** : combler ce gap **sans** installer l'outil tiers (860 KB de Python, mainteneur
|
||||||
|
unique, surface supply-chain + credentials OAuth à confier). On ajoute une couche data
|
||||||
|
minimale, isolée, sous contrôle, branchée sur les analyzers existants.
|
||||||
|
|
||||||
|
**Ce que ça débloque concrètement** :
|
||||||
|
- CWV **terrain** (CrUX, 75e percentile, mobile + desktop, historique) au lieu du seul labo.
|
||||||
|
- Requêtes GSC : le pattern « **positions 4-10 à fort volume d'impressions** » = quick wins
|
||||||
|
que les skills ne pouvaient pas voir.
|
||||||
|
- Indexation réelle par URL (GSC URL Inspection) au lieu d'une déduction.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Principes directeurs (non négociables)
|
||||||
|
|
||||||
|
1. **Sécurité avant tout.** Secrets hors git, permissions `0600`, scope OAuth **lecture seule**,
|
||||||
|
redaction systématique, aucun secret dans un rapport/log. Toute surface secret nouvelle sous
|
||||||
|
`~/.claude` est **explicitement allowlistée** dans `.gitleaks.toml`.
|
||||||
|
2. **Dégradation gracieuse (fail-open audit).** Creds absents / token révoqué / quota 429 →
|
||||||
|
l'audit FULL **continue** en retombant sur PageSpeed anonyme + une action utilisateur.
|
||||||
|
Jamais de crash.
|
||||||
|
3. **Consentement OAuth one-shot, runs silencieux ensuite.** Le consentement navigateur ne se
|
||||||
|
fait qu'au setup (ou à l'ajout d'un compte). Les audits suivants sont non-interactifs.
|
||||||
|
4. **Multi-compte, zéro conflit.** Plusieurs comptes Google connectables ; deux audits de deux
|
||||||
|
sites en simultané sont **isolés par construction** (compte + propriété = paramètres explicites
|
||||||
|
par appel, jamais un état global mutable).
|
||||||
|
5. **Isolation.** Le moteur ne connaît rien du SEO (rend du JSON) ; les analyzers ne connaissent
|
||||||
|
rien d'OAuth (consomment du JSON). Deps Python isolées dans un venv dédié.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Décisions actées
|
||||||
|
|
||||||
|
| # | Décision | Choix |
|
||||||
|
|---|---|---|
|
||||||
|
| Auth GSC | OAuth2 installed-app, consentement one-time, refresh token stocké | **OAuth2** |
|
||||||
|
| Scope data v1 | GSC Search Analytics + URL Inspection + CrUX field | **oui** (GA4/Ads/Indexing hors v1) |
|
||||||
|
| Surface | Fold dans `/seo` (+`/geo`) FULL ; pas de nouveau skill d'audit | **oui** (setup = `make seo-connect`, pas un skill d'audit) |
|
||||||
|
| Langage | Helper Python + `google-auth`, venv isolé | **oui** |
|
||||||
|
| Multi-compte | Store keyé par compte ; sélection à **chaque** audit FULL | **oui** |
|
||||||
|
| Persistance token | Auto-écriture idempotente dans le store (write-temp→rename) | **oui** |
|
||||||
|
| Déclencheur consentement | `install.sh` (proposé) **et** `make seo-connect` (toujours dispo) | **oui, les deux** |
|
||||||
|
| Gitleaks | Allowlister le token store (comme `~/.claude/.env`) | **oui** |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Architecture
|
||||||
|
|
||||||
|
### 4.1 Vue d'ensemble
|
||||||
|
|
||||||
|
```
|
||||||
|
lib/seo-data/ ← LE MOTEUR (nouveau, isolé, sans logique SEO)
|
||||||
|
├── fetch.sh entrypoint bash : source ~/.claude/.env → active venv → dispatch
|
||||||
|
│ sous-commandes → JSON sur stdout → dégrade proprement si creds absents
|
||||||
|
├── google_seo.py appels GSC (Search Analytics, URL Inspection, sites.list) + CrUX,
|
||||||
|
│ refresh OAuth via google-auth, normalisation JSON
|
||||||
|
├── connect.py consentement OAuth one-time (InstalledAppFlow) + écriture store
|
||||||
|
├── tokenstore.py lecture/écriture atomique du store keyé (partagé par connect+fetch)
|
||||||
|
├── requirements.txt deps épinglées : google-auth, google-auth-oauthlib, requests
|
||||||
|
└── README.md contrat d'usage + sous-commandes
|
||||||
|
|
||||||
|
~/.claude/.venv-seo-data/ ← venv isolé (deps hors système, reproductible)
|
||||||
|
~/.claude/seo-data/tokens.json ← store keyé par compte (0600, hors git)
|
||||||
|
|
||||||
|
CONSOMMATEURS (existants, patchés) :
|
||||||
|
agents/seo-analyzer.md STEP 4 (CWV) → data terrain CrUX quand dispo ; nouvelle
|
||||||
|
sous-section « Performance GSC » ; STEP 9 axe Technical nourri au réel.
|
||||||
|
skills/seo/SKILL.md STEP 0 → sélection compte + propriété (main loop, interactif).
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.2 Composants & frontières
|
||||||
|
|
||||||
|
| Unité | Rôle unique | Utilisée comment | Dépend de |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `fetch.sh` | orchestre env→venv→python, dégrade, redige | `bash fetch.sh <cmd> --account … --property …` → JSON | `~/.claude/.env`, venv, tokenstore |
|
||||||
|
| `google_seo.py` | appelle GSC + CrUX, normalise en JSON | appelé par `fetch.sh` | google-auth, requests |
|
||||||
|
| `connect.py` | consentement OAuth one-time + découverte propriétés | `make seo-connect` / STEP 0 | google-auth-oauthlib, tokenstore |
|
||||||
|
| `tokenstore.py` | I/O atomique du store keyé | importé par connect + google_seo | stdlib (json, os, fcntl) |
|
||||||
|
| seo-analyzer (patch) | consomme le JSON, score, rapporte | inchangé pour l'utilisateur | `fetch.sh` (optionnel) |
|
||||||
|
| seo/SKILL.md (patch) | sélectionne compte+propriété en STEP 0 | interactif, main loop | `fetch.sh accounts` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Authentification & secrets (multi-compte)
|
||||||
|
|
||||||
|
### 5.1 Modèle OAuth
|
||||||
|
|
||||||
|
- **Type** : OAuth2 « installed app » (client Desktop créé dans la console GCP par l'utilisateur).
|
||||||
|
- **Scope unique** : `https://www.googleapis.com/auth/webmasters.readonly` (GSC lecture seule).
|
||||||
|
Least privilege strict : le token ne peut **rien modifier** sur GSC, révocable côté Google.
|
||||||
|
- **Flow** : `connect.py` construit la config client **en mémoire** (`InstalledAppFlow.from_client_config`)
|
||||||
|
à partir des vars d'env — **aucun `client_secret.json` sur disque**. Le consentement ouvre un
|
||||||
|
serveur local + navigateur ; au retour, on obtient un `refresh_token` (durable) écrit dans le store.
|
||||||
|
- **Runs d'audit** : `google_seo.py` échange le `refresh_token` contre un `access_token` **éphémère
|
||||||
|
en mémoire** (jamais persisté). Donc **aucune écriture disque pendant un audit**.
|
||||||
|
|
||||||
|
### 5.2 Vault `~/.claude/.env` (app partagée + CrUX)
|
||||||
|
|
||||||
|
Ne contient que ce qui est **commun à tous les comptes** :
|
||||||
|
|
||||||
|
```
|
||||||
|
# ── Google SEO data layer (lib/seo-data) ──
|
||||||
|
# App OAuth Desktop partagée (console GCP → APIs & Services → Identifiants).
|
||||||
|
# Scope demandé : webmasters.readonly. Setup : make seo-connect
|
||||||
|
GOOGLE_OAUTH_CLIENT_ID=<votre-client-id.apps.googleusercontent.com>
|
||||||
|
GOOGLE_OAUTH_CLIENT_SECRET=<votre-client-secret>
|
||||||
|
# Clé API CrUX + PageSpeed (console GCP → clé API restreinte à ces deux APIs).
|
||||||
|
# Get it: https://developer.chrome.com/docs/crux/api (bouton "Get a key")
|
||||||
|
CRUX_API_KEY=<votre-cle-crux>
|
||||||
|
```
|
||||||
|
|
||||||
|
Les **refresh tokens ne sont PAS ici** (multi-compte → store dédié).
|
||||||
|
|
||||||
|
### 5.3 Token store keyé
|
||||||
|
|
||||||
|
`~/.claude/seo-data/tokens.json`, permissions `0600`, dossier `0700` :
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": 1,
|
||||||
|
"accounts": {
|
||||||
|
"client-a": {
|
||||||
|
"refresh_token": "<opaque>",
|
||||||
|
"scopes": ["https://www.googleapis.com/auth/webmasters.readonly"],
|
||||||
|
"granted_at": "2026-07-09T…",
|
||||||
|
"properties": ["sc-domain:site-a.com", "https://www.site-a.com/"]
|
||||||
|
},
|
||||||
|
"client-b": { "…": "…" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Clé = **label choisi par l'utilisateur** au moment du `connect` (ex. `client-a`), **pas** l'email.
|
||||||
|
Raison sécurité : keyer par email obligerait à élargir le scope OAuth (`userinfo.email`) juste pour
|
||||||
|
l'identification. On reste à `webmasters.readonly` strict ; le label suffit à distinguer les comptes.
|
||||||
|
Collision de label → `connect` demande confirmation (écraser / renommer).
|
||||||
|
- `properties` = propriétés GSC accessibles (via `sites.list`, **déjà** dans le scope
|
||||||
|
`webmasters.readonly`), pré-remplies au `connect` pour la sélection en STEP 0.
|
||||||
|
- Écriture **uniquement** au `connect` (jamais pendant un audit), **atomique** : write vers
|
||||||
|
`tokens.json.tmp` → `fsync` → `rename` ; verrou `fcntl` exclusif le temps de l'échange
|
||||||
|
read-modify-write pour couvrir deux `connect` simultanés.
|
||||||
|
|
||||||
|
### 5.4 Gitleaks allowlist + gitignore
|
||||||
|
|
||||||
|
- `~/.claude/seo-data/tokens.json` vit **hors du repo** (le repo ne symlink que
|
||||||
|
`hooks agents skills lib templates rules`). Il n'entre donc jamais en git directement.
|
||||||
|
- Mais `make scan-secrets` (gitleaks) balaie `~/.claude` à la recherche de copies de secrets.
|
||||||
|
On **ajoute une entrée d'allowlist** dans `.gitleaks.toml` `[allowlist].paths`, exactement
|
||||||
|
comme `~/.claude/.env` l'est déjà :
|
||||||
|
```toml
|
||||||
|
# Token store OAuth de la couche seo-data — secret local légitime (BDR-026 pattern),
|
||||||
|
# hors git, 0600. On l'allowliste pour ne pas noyer scan-secrets de faux positifs.
|
||||||
|
'''(^|/)\.claude/seo-data/tokens\.json$''',
|
||||||
|
```
|
||||||
|
- `.gitignore` : ajouter `.venv-seo-data/` et `seo-data/tokens.json` par prudence (au cas où un
|
||||||
|
chemin relatif les ferait apparaître sous le repo), en complément de l'exclusion `.env*` existante.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Sélection de compte & propriété (STEP 0, main loop)
|
||||||
|
|
||||||
|
Interactif → se déroule **dans le dispatcher `/seo` (main loop)**, jamais dans le subagent
|
||||||
|
(qui ne peut pas interagir). Uniquement en **FULL** (LOCAL n'a pas de donnée live).
|
||||||
|
|
||||||
|
1. Lister les comptes connectés : `bash lib/seo-data/fetch.sh accounts` → JSON `{accounts:[…]}`.
|
||||||
|
2. Présenter à l'utilisateur (label + propriétés découvertes) :
|
||||||
|
```
|
||||||
|
COMPTE GOOGLE pour cet audit FULL :
|
||||||
|
1) client-a (sc-domain:site-a.com, https://www.site-a.com/)
|
||||||
|
2) client-b (sc-domain:site-b.com)
|
||||||
|
N) Connecter un nouveau compte (choisir un label)
|
||||||
|
S) Ignorer (audit sans donnée GSC — CWV terrain via CrUX seulement)
|
||||||
|
```
|
||||||
|
3. « Connecter un nouveau compte » → demande un **label**, lance `connect.py` dans le main loop
|
||||||
|
(consentement navigateur), puis auto-découverte des propriétés (`sites.list`) → re-liste.
|
||||||
|
4. Compte choisi → si plusieurs propriétés, demander **laquelle** correspond au site audité.
|
||||||
|
5. Le couple `(account, property)` retenu est passé **explicitement** dans le contexte de
|
||||||
|
l'analyzer (bloc BUSINESS CONTEXT du dispatch, STEP 1), qui appellera
|
||||||
|
`fetch.sh queries|inspect --account <a> --property <p>`.
|
||||||
|
|
||||||
|
CrUX ne demande pas de compte (clé API publique) → toujours tenté si `CRUX_API_KEY` présent,
|
||||||
|
indépendamment du choix de compte.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Sûreté concurrentielle (2 sites en parallèle)
|
||||||
|
|
||||||
|
Garantie **par construction**, pas par verrou global :
|
||||||
|
|
||||||
|
- **Pas d'état « compte courant ».** Le compte + la propriété sont des **arguments explicites**
|
||||||
|
de chaque `fetch.sh`. Deux audits (2 sessions Claude, ou 2 sites) ne partagent aucune variable
|
||||||
|
mutable de sélection.
|
||||||
|
- **Audits = lecture seule** du store. Les access tokens sont éphémères en mémoire, jamais écrits.
|
||||||
|
Donc deux audits concurrents ne s'écrivent jamais dessus.
|
||||||
|
- **Écriture = seulement au `connect`**, atomique (`tmp`→`fsync`→`rename`) sous verrou `fcntl`,
|
||||||
|
pour couvrir le cas rare de deux consentements simultanés.
|
||||||
|
- Le venv est en lecture seule à l'exécution (créé/maj uniquement par `make seo-connect`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Périmètre data & mapping dans le rapport
|
||||||
|
|
||||||
|
| Donnée | Source | Sous-commande | Atterrit dans `SEO.md` |
|
||||||
|
|---|---|---|---|
|
||||||
|
| CWV terrain (LCP/INP/CLS 75e pct, mobile+desktop, historique) | CrUX API | `fetch.sh crux` | §2 Audit technique — **note primaire** ; PageSpeed labo gardé en secondaire diagnostic |
|
||||||
|
| Requêtes (impressions, clics, CTR, position) | GSC Search Analytics | `fetch.sh queries` | §2/§8 — sous-section « Performance GSC » + **quick wins position 4-10** |
|
||||||
|
| Pages (perf par URL) | GSC Search Analytics | `fetch.sh queries --dim page` | idem — top pages |
|
||||||
|
| Indexation par URL | GSC URL Inspection | `fetch.sh inspect` | §2 indexabilité — **fait** vs déduction |
|
||||||
|
|
||||||
|
- **Scoring** : l'axe *Technical* (STEP 9 de `seo-analyzer`) se calcule sur le **terrain** quand
|
||||||
|
dispo ; sinon labo (dégradation).
|
||||||
|
- **Nouveau contenu de rapport** : une sous-section « Performance GSC (90 j) » dans §2, listant top
|
||||||
|
requêtes + les quick wins position 4-10. Reste en **français**, cohérent avec l'existant.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Interface du moteur (`fetch.sh`)
|
||||||
|
|
||||||
|
Contrat stable que les analyzers consomment (JSON sur stdout, exit 0 même en dégradé) :
|
||||||
|
|
||||||
|
```bash
|
||||||
|
fetch.sh accounts
|
||||||
|
→ {"status":"ok","accounts":[{"label":"…","properties":[…],"granted_at":"…"}]} # [] si aucun compte
|
||||||
|
|
||||||
|
fetch.sh crux --url https://ex.com [--strategy mobile|desktop]
|
||||||
|
→ {"status":"ok","source":"crux","lcp_p75_ms":…,"inp_p75_ms":…,"cls_p75":…} # métrique absente = clé omise
|
||||||
|
→ {"status":"degraded","reason":"no_crux_key"|"no_field_data"|"rate_limited"}
|
||||||
|
# 404 page-level → retry automatique origin-level avant de dégrader
|
||||||
|
|
||||||
|
fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--dim query|page]
|
||||||
|
→ {"status":"ok","source":"gsc","rows":[{"key":"…","clicks":…,"impressions":…,"ctr":…,"position":…}]}
|
||||||
|
→ {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"}
|
||||||
|
|
||||||
|
fetch.sh inspect --account client-a --property … --url https://ex.com/page
|
||||||
|
→ {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…"}
|
||||||
|
→ {"status":"degraded","reason":"…"}
|
||||||
|
```
|
||||||
|
|
||||||
|
Règles : **jamais** de secret dans la sortie ; messages d'erreur génériques ; exit 0 en dégradé
|
||||||
|
(l'analyzer décide de la suite), exit ≠ 0 uniquement sur mauvais usage (args invalides, exit 2).
|
||||||
|
**Toute erreur imprévue** (HTTP 403/5xx, timeout, DNS) → `{"status":"degraded","reason":"unexpected_error"}`
|
||||||
|
+ exit 0 — jamais de traceback, jamais de stdout vide. Tests : `SEO_DATA_ENV_FILE` permet de
|
||||||
|
substituer le vault (les tests pointent `/dev/null` — jamais le vrai `~/.claude/.env`) ;
|
||||||
|
`SEO_DATA_DEBUG=1` réactive stderr pour diagnostiquer.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Dégradation gracieuse & posture sécurité
|
||||||
|
|
||||||
|
- **Fail-open audit / fail-closed data** : creds manquants, refresh échoué, 429 → `{"status":"degraded"}`,
|
||||||
|
exit 0. L'analyzer bascule sur PageSpeed anonyme et émet en §11 « Connecter GSC : `make seo-connect` »
|
||||||
|
(réutilise la formulation `automation-catalog.md`).
|
||||||
|
- **Redaction** : `fetch.sh` ne logge jamais les variables d'env ni le token ; stdout = JSON de
|
||||||
|
données uniquement ; stderr = messages génériques.
|
||||||
|
- **Least privilege** : scope `webmasters.readonly` seul ; clé CrUX restreinte à CrUX + PageSpeed.
|
||||||
|
- **Supply-chain maîtrisée** : 3 libs Google officielles, **épinglées** dans `requirements.txt`,
|
||||||
|
isolées dans un venv — surface auditablement listée, sans commune mesure avec l'outil tiers.
|
||||||
|
- **Reprise sur token révoqué** : `doctor.sh` signale, message pointe vers `make seo-connect` pour re-consentir.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Install & déploiement (touch-list)
|
||||||
|
|
||||||
|
Séquence respectant l'ordre critique **`link.sh` → vault joignable → consentement** (évite le
|
||||||
|
blocker connu : le symlink `~/.claude/.env` est créé par `link.sh`, absent sur machine fraîche ;
|
||||||
|
tout lecteur de creds vise le **canonical `~/.claude/.env`**, pas `$REPO/.env`).
|
||||||
|
|
||||||
|
| Fichier | Modification |
|
||||||
|
|---|---|
|
||||||
|
| `.env.example` | Ajouter les 3 vars (client id/secret, CrUX key) au format existant (`# Used by:` / `# Get it:` + placeholder). **Seul fichier versionné touché côté secrets.** |
|
||||||
|
| `install.sh` | Après `link.sh` (§5) et `claude login` (§3) : étape **optionnelle idempotente** (moule « Press Enter to connect… ») → si aucun compte dans le store, propose `make seo-connect` ; skip sinon. |
|
||||||
|
| `Makefile` | Cible user-facing `seo-connect` (venv + `pip install -r lib/seo-data/requirements.txt` + `connect.py`, rejouable) **et** extension de la cible `test` pour découvrir `lib/seo-data/*.test.sh` (le test vit hors `lib/tests/`, gaté par `config-protection.sh`). |
|
||||||
|
| `doctor.sh` | Nouveau check (lit `~/.claude/.env` canonical) : venv + deps présents ? au moins un compte dans le store ? `CRUX_API_KEY` présent ? → **PASS / WARN, jamais fatal**. |
|
||||||
|
| `.gitleaks.toml` | Allowlist du token store (cf. §5.4). |
|
||||||
|
| `.gitignore` | Ajouter `.venv-seo-data/` et `seo-data/tokens.json` (ceinture + bretelles). |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Tests
|
||||||
|
|
||||||
|
`lib/seo-data/seo-data.test.sh`, convention du repo (`tf`/`tr_`/`tn` + compteurs PASS/FAIL),
|
||||||
|
**sans appel réseau**. Placé sous `lib/seo-data/` (co-localisé, **hors `lib/tests/`** que
|
||||||
|
`hooks/config-protection.sh` protège comme dossier-garde) ; `make test` le découvre via le glob
|
||||||
|
ajouté en Task 6. En TDD : `bash lib/seo-data/seo-data.test.sh`.
|
||||||
|
|
||||||
|
- Parsing des sous-commandes/args de `fetch.sh` (bons/mauvais usages, exit codes).
|
||||||
|
- Parsing de forme JSON sur **fixtures commitées** (réponses GSC/CrUX mockées) → shape attendue.
|
||||||
|
- **Redaction** : erreur simulée (token bidon) → aucune valeur secrète dans stdout/stderr.
|
||||||
|
- **Dégradation** : creds absents → `{"status":"degraded"}` + **exit 0** (pas 1).
|
||||||
|
- Isolement : deux invocations `--account` différentes → sélections indépendantes (pas d'état partagé).
|
||||||
|
|
||||||
|
Fixtures sous `lib/seo-data/fixtures/` (réponses synthétiques, aucun vrai secret/PII).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Hors périmètre (YAGNI v1)
|
||||||
|
|
||||||
|
- GA4 (trafic organique), Google Ads / Keyword Planner, Indexing API.
|
||||||
|
- Modification de `geo-analyzer` (le GSC pourrait plus tard éclairer quelles requêtes déclenchent
|
||||||
|
des AI Overviews — v2).
|
||||||
|
- Audit multi-propriétés en un seul run (une propriété par audit en v1).
|
||||||
|
- Monitoring/drift dans le temps (SQLite) — l'outil tiers le fait ; hors scope v1.
|
||||||
|
- CrUX History API (tendance 25 semaines) — v1 = snapshot p75 seulement.
|
||||||
|
- Routage de l'appel PageSpeed via `fetch.sh` avec `CRUX_API_KEY` (dé-quota) — **rejeté en v1
|
||||||
|
pour raison sécurité** : passer la clé à l'analyzer l'exposerait dans le contexte du subagent
|
||||||
|
(ligne de commande curl → risque de fuite dans rapport/log). v2 : sous-commande
|
||||||
|
`fetch.sh pagespeed` où la clé reste confinée au moteur.
|
||||||
|
- Nouveau skill d'audit dédié `/gsc` (le setup passe par `make seo-connect`, l'usage par `/seo` FULL).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. Liste des fichiers touchés
|
||||||
|
|
||||||
|
**Créés**
|
||||||
|
- `lib/seo-data/fetch.sh`
|
||||||
|
- `lib/seo-data/google_seo.py`
|
||||||
|
- `lib/seo-data/connect.py`
|
||||||
|
- `lib/seo-data/tokenstore.py`
|
||||||
|
- `lib/seo-data/requirements.txt`
|
||||||
|
- `lib/seo-data/README.md`
|
||||||
|
- `lib/seo-data/seo-data.test.sh`
|
||||||
|
- `lib/seo-data/fixtures/*.json`
|
||||||
|
|
||||||
|
**Modifiés**
|
||||||
|
- `.env.example` (3 vars)
|
||||||
|
- `install.sh` (étape consentement optionnelle post-link)
|
||||||
|
- `Makefile` (cible `seo-connect`)
|
||||||
|
- `doctor.sh` (check creds non-fatal)
|
||||||
|
- `.gitleaks.toml` (allowlist token store)
|
||||||
|
- `.gitignore` (venv + token store)
|
||||||
|
- `agents/seo-analyzer.md` (STEP 4 CWV terrain + sous-section Performance GSC ; STEP 9 axe Technical)
|
||||||
|
- `skills/seo/SKILL.md` (STEP 0 sélection compte + propriété en FULL)
|
||||||
|
- `agents/resources/automation-catalog.md` (section « Google Search Console — connexion OAuth » réutilisable en §11)
|
||||||
|
|
||||||
|
**Hors repo (générés au setup, jamais commités)**
|
||||||
|
- `~/.claude/.venv-seo-data/`
|
||||||
|
- `~/.claude/seo-data/tokens.json`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 15. Risques & mitigations
|
||||||
|
|
||||||
|
| Risque | Mitigation |
|
||||||
|
|---|---|
|
||||||
|
| Symlink `~/.claude/.env` absent sur machine fraîche | Lecture du **canonical** `~/.claude/.env` ; consentement **après** `link.sh` |
|
||||||
|
| Deux `connect` simultanés corrompent le store | Écriture atomique `tmp`→`rename` + verrou `fcntl` |
|
||||||
|
| Deux audits sur 2 sites se marchent dessus | Compte+propriété **explicites par appel** ; audits en lecture seule |
|
||||||
|
| Secret loggé par erreur | Redaction imposée + test dédié |
|
||||||
|
| Token store flaggé par scan-secrets | Allowlist gitleaks explicite (§5.4) |
|
||||||
|
| `python3`/venv absent | `doctor.sh` warn ; `make seo-connect` crée le venv ; dégradation si absent |
|
||||||
|
| Quota GSC/CrUX (429) | Traité comme `degraded` → audit continue |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 16. Questions ouvertes résolues
|
||||||
|
|
||||||
|
- **Scopes** → `webmasters.readonly` seul.
|
||||||
|
- **CrUX vs PageSpeed** → les deux : CrUX terrain primaire, PageSpeed labo secondaire/fallback.
|
||||||
|
- **Forme data en §2** → terrain 75e pct primaire, labo en secondaire.
|
||||||
|
- **GSC obligatoire ?** → non, optionnel, dégradation gracieuse ; proposé à chaque FULL.
|
||||||
|
- **Token expiré** → `degraded` + `doctor.sh` pointe `make seo-connect`.
|
||||||
|
- **Locale** → rapports en français, cohérent avec l'existant.
|
||||||
|
- **Multi-compte** → store keyé par **label utilisateur** (scope inchangé) ; sélection par audit ;
|
||||||
|
isolation par arguments explicites.
|
||||||
|
```
|
||||||
@@ -400,6 +400,21 @@ fi
|
|||||||
|
|
||||||
echo ""
|
echo ""
|
||||||
|
|
||||||
|
# ── seo-data (GSC/CrUX data layer) — non-fatal ──
|
||||||
|
ENVF="$HOME/.claude/.env"
|
||||||
|
if grep -qE '^[[:space:]]*(export[[:space:]]+)?CRUX_API_KEY=.' "$ENVF" 2>/dev/null; then
|
||||||
|
pass "seo-data: CRUX_API_KEY present"
|
||||||
|
else
|
||||||
|
warn "seo-data: CRUX_API_KEY absent in ~/.claude/.env — /seo FULL falls back to lab PageSpeed"
|
||||||
|
fi
|
||||||
|
STORE="$HOME/.claude/seo-data/tokens.json"
|
||||||
|
if [ -f "$STORE" ]; then
|
||||||
|
N=$(python3 "$REPO/lib/seo-data/tokenstore.py" list --file "$STORE" 2>/dev/null | grep -o '"label"' | wc -l)
|
||||||
|
pass "seo-data: $N Google account(s) connected"
|
||||||
|
else
|
||||||
|
warn "seo-data: no Google account connected (run: make seo-connect) — GSC data disabled"
|
||||||
|
fi
|
||||||
|
|
||||||
# ────────────────────────────────────────────────────────────
|
# ────────────────────────────────────────────────────────────
|
||||||
# Summary
|
# Summary
|
||||||
# ────────────────────────────────────────────────────────────
|
# ────────────────────────────────────────────────────────────
|
||||||
|
|||||||
+10
@@ -106,6 +106,16 @@ echo ""
|
|||||||
echo "── Setting up symlinks..."
|
echo "── Setting up symlinks..."
|
||||||
bash "$REPO/link.sh"
|
bash "$REPO/link.sh"
|
||||||
|
|
||||||
|
# ── 5b. Optional: connect a Google account for /seo FULL ──
|
||||||
|
echo ""
|
||||||
|
if [ -f "$HOME/.claude/seo-data/tokens.json" ]; then
|
||||||
|
ok "seo-data: a Google account is already connected"
|
||||||
|
else
|
||||||
|
info "SEO data layer (GSC + CrUX) is optional. To enable real Search Console"
|
||||||
|
info "data in /seo FULL: add GOOGLE_OAUTH_* + CRUX_API_KEY to ~/.claude/.env,"
|
||||||
|
info "then run: make seo-connect"
|
||||||
|
fi
|
||||||
|
|
||||||
# ── 6. Install plugins ──
|
# ── 6. Install plugins ──
|
||||||
echo ""
|
echo ""
|
||||||
echo "── Installing plugins..."
|
echo "── Installing plugins..."
|
||||||
|
|||||||
@@ -0,0 +1,192 @@
|
|||||||
|
# seo-data — GSC + CrUX data layer for `/seo` FULL audits
|
||||||
|
|
||||||
|
Small, isolated engine that gives the `/seo` skill real Google data instead of
|
||||||
|
guesses: **Search Console** (queries, positions, indexation) and **CrUX**
|
||||||
|
(Core Web Vitals *field* data — real users, not lab simulation). It knows
|
||||||
|
nothing about SEO scoring; it only turns Google APIs into normalized JSON.
|
||||||
|
The `seo-analyzer` agent consumes that JSON in STEP 4 (Core Web Vitals) and
|
||||||
|
the new "Performance GSC" subsection; the `/seo` skill selects the account
|
||||||
|
and property in STEP 0 of a FULL audit (not needed for LOCAL).
|
||||||
|
|
||||||
|
Multi-account by design: the token store is keyed by a user-chosen label, and
|
||||||
|
every call takes `--account`/`--property` explicitly. Two audits running at
|
||||||
|
the same time (two sites, two sessions) never share mutable state — nothing
|
||||||
|
is written to disk during an audit, only at `make seo-connect`.
|
||||||
|
|
||||||
|
## Setup
|
||||||
|
|
||||||
|
One-time per Google account:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make seo-connect
|
||||||
|
```
|
||||||
|
|
||||||
|
This creates `~/.claude/.venv-seo-data/` (isolated venv, deps pinned in
|
||||||
|
`requirements.txt`), installs `google-auth`, `google-auth-oauthlib`,
|
||||||
|
`requests`, then runs `connect.py`: it opens a browser for OAuth consent and
|
||||||
|
asks for a **label** (e.g. `client-a`) to key the account — pick a name, not
|
||||||
|
an email, since the store never stores or requests the account's email.
|
||||||
|
|
||||||
|
Before running it, set these 3 keys in `~/.claude/.env` (the canonical
|
||||||
|
vault; `link.sh` only symlinks the repo's `.env` to it and warns with a
|
||||||
|
`cp .env.example .env` hint if it's missing — it never creates the vault
|
||||||
|
itself):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
GOOGLE_OAUTH_CLIENT_ID=<your-client-id>.apps.googleusercontent.com
|
||||||
|
GOOGLE_OAUTH_CLIENT_SECRET=<your-client-secret>
|
||||||
|
CRUX_API_KEY=<your-crux-api-key>
|
||||||
|
```
|
||||||
|
|
||||||
|
- `GOOGLE_OAUTH_CLIENT_ID` / `GOOGLE_OAUTH_CLIENT_SECRET` — OAuth2 "Desktop
|
||||||
|
app" credentials from the Google Cloud Console (APIs & Services →
|
||||||
|
Credentials). Shared across every account you connect; the OAuth scope
|
||||||
|
requested is `https://www.googleapis.com/auth/webmasters.readonly` only
|
||||||
|
— read-only Search Console, nothing can be modified or deleted via this
|
||||||
|
token.
|
||||||
|
- `CRUX_API_KEY` — a Chrome UX Report API key (restrict it to CrUX +
|
||||||
|
PageSpeed in the Console). Get one at
|
||||||
|
https://developer.chrome.com/docs/crux/api. No OAuth involved: CrUX is
|
||||||
|
public field data, gated by API key only, independent of any connected
|
||||||
|
account.
|
||||||
|
|
||||||
|
`make seo-connect` is idempotent and rerunnable — connecting a second
|
||||||
|
account just runs it again with a different label; reusing an existing
|
||||||
|
label prompts to overwrite.
|
||||||
|
|
||||||
|
## `fetch.sh` contract
|
||||||
|
|
||||||
|
`lib/seo-data/fetch.sh` is the one stable entrypoint analyzers call. It
|
||||||
|
sources `~/.claude/.env`, prefers the isolated venv (falls back to system
|
||||||
|
`python3` for stdlib-only paths), dispatches to `google_seo.py` or
|
||||||
|
`tokenstore.py`, and never prints a secret to stdout or stderr.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
fetch.sh accounts
|
||||||
|
→ {"status":"ok","accounts":[{"label":"…","properties":[…],"granted_at":"…"}]} # [] if none connected
|
||||||
|
|
||||||
|
fetch.sh crux --url https://ex.com [--strategy mobile|desktop]
|
||||||
|
→ {"status":"ok","source":"crux","lcp_p75_ms":…,"inp_p75_ms":…,"cls_p75":…} # a missing metric omits its key
|
||||||
|
→ {"status":"degraded","reason":"no_crux_key"|"no_field_data"|"rate_limited"}
|
||||||
|
# a 404 on page-level data retries at origin-level before degrading
|
||||||
|
|
||||||
|
fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--dim query|page]
|
||||||
|
→ {"status":"ok","source":"gsc","dimension":"query","rows":[{"key":"…","clicks":…,"impressions":…,"ctr":…,"position":…}]}
|
||||||
|
→ {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"}
|
||||||
|
|
||||||
|
fetch.sh inspect --account client-a --property … --url https://ex.com/page
|
||||||
|
→ {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…"}
|
||||||
|
→ {"status":"degraded","reason":"…"}
|
||||||
|
```
|
||||||
|
|
||||||
|
Rules that hold for every subcommand:
|
||||||
|
|
||||||
|
- **JSON always on stdout, never empty.** Even an unexpected error (HTTP
|
||||||
|
403/5xx, timeout, DNS failure) prints
|
||||||
|
`{"status":"degraded","reason":"unexpected_error"}` — never a raw
|
||||||
|
traceback.
|
||||||
|
- **`status` is `"ok"` or `"degraded"` on exit 0; `"error"` on exit 2.**
|
||||||
|
Analyzers branch on this field; `"error"` only shows up on bad usage,
|
||||||
|
`reason` is informational otherwise.
|
||||||
|
- **Exit code 0 on `ok` and on `degraded`.** The engine never fails the
|
||||||
|
process just because Google data isn't available — that's a normal,
|
||||||
|
expected outcome the analyzer handles by falling back. **Exit code 2**
|
||||||
|
is reserved for bad usage: unknown subcommand, missing required flag,
|
||||||
|
invalid argument — those paths emit `{"status":"error",...}` instead.
|
||||||
|
- **`--store` is accepted uniformly** by every subcommand for consistent
|
||||||
|
`fetch.sh` dispatch, even though `crux` ignores it (CrUX needs no
|
||||||
|
account).
|
||||||
|
- **Never prints a secret.** No env var, refresh token, or access token
|
||||||
|
ever reaches stdout or stderr, including in error paths.
|
||||||
|
|
||||||
|
Two env vars exist for testing, never for normal use:
|
||||||
|
`SEO_DATA_ENV_FILE` overrides which env file is sourced (tests point it at
|
||||||
|
`/dev/null` so a real `~/.claude/.env` on the machine can never leak into a
|
||||||
|
test run), and `SEO_DATA_DEBUG=1` re-enables stderr for local debugging
|
||||||
|
(stderr is suppressed by default so library warnings can't leak a secret
|
||||||
|
into an agent's context).
|
||||||
|
|
||||||
|
## Token store
|
||||||
|
|
||||||
|
`~/.claude/seo-data/tokens.json` — refresh tokens, keyed by the label chosen
|
||||||
|
at `make seo-connect`, one entry per connected account:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": 1,
|
||||||
|
"accounts": {
|
||||||
|
"client-a": {
|
||||||
|
"refresh_token": "<opaque>",
|
||||||
|
"scopes": ["https://www.googleapis.com/auth/webmasters.readonly"],
|
||||||
|
"granted_at": "2026-07-09T12:00:00+00:00",
|
||||||
|
"properties": ["sc-domain:site-a.com", "https://www.site-a.com/"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Security posture:
|
||||||
|
|
||||||
|
- **File `0600`, directory `0700`.** `tokenstore.save_account` re-asserts
|
||||||
|
both permissions on every write.
|
||||||
|
- **Written only at `connect` time, atomically.** `tmp` → `fsync` →
|
||||||
|
`os.replace` (atomic rename), under an exclusive `fcntl` lock, so two
|
||||||
|
simultaneous `make seo-connect` runs can't corrupt the file. Audits never
|
||||||
|
write to this file — access tokens are exchanged in memory and never
|
||||||
|
persisted, so two audits running concurrently never contend on it.
|
||||||
|
- **Keyed by label, not email.** Identifying accounts by email would
|
||||||
|
require widening the OAuth scope just for identification; the label the
|
||||||
|
user picks at connect time is sufficient and keeps the scope at
|
||||||
|
`webmasters.readonly` only (least privilege).
|
||||||
|
- **Refresh tokens are redacted from `list`.** `fetch.sh accounts` (and
|
||||||
|
`tokenstore.py list`) return label, properties, and `granted_at` only —
|
||||||
|
the `refresh_token` field is intentionally never included in that output.
|
||||||
|
- **Allowlisted in gitleaks.** The store lives under `~/.claude/`, outside
|
||||||
|
this repo, so it's never committed directly — but `make scan-secrets`
|
||||||
|
also sweeps `~/.claude` for stray copies of secrets. `.gitleaks.toml` has
|
||||||
|
an explicit `[allowlist].paths` entry for
|
||||||
|
`(^|/)\.claude/seo-data/tokens\.json$`, the same treatment
|
||||||
|
`~/.claude/.env` already gets, so a legitimate local secret store doesn't
|
||||||
|
drown real findings in false positives.
|
||||||
|
- **Also gitignored** (`.venv-seo-data/` and `seo-data/tokens.json` in
|
||||||
|
`.gitignore`) as a second, belt-and-suspenders guard in case a relative
|
||||||
|
path ever put either under the repo tree.
|
||||||
|
|
||||||
|
## Graceful degradation
|
||||||
|
|
||||||
|
Missing API key, no connected account, or a revoked/expired token is a
|
||||||
|
**normal outcome, not a failure**:
|
||||||
|
|
||||||
|
- No `CRUX_API_KEY` → `crux` returns `{"status":"degraded","reason":"no_crux_key"}`.
|
||||||
|
- No account connected, or the store has no refresh token for the given
|
||||||
|
`--account` → `queries`/`inspect` return
|
||||||
|
`{"status":"degraded","reason":"no_credentials"}`.
|
||||||
|
- Refresh token revoked at Google's end → `{"status":"degraded","reason":"token_revoked"}`
|
||||||
|
(a transient network blip during refresh is classified
|
||||||
|
`"network_error"` instead, so a flaky connection never forces the user
|
||||||
|
back through OAuth).
|
||||||
|
- Rate limited (HTTP 429) on any Google API → `{"status":"degraded","reason":"rate_limited"}`.
|
||||||
|
|
||||||
|
In every case: **exit code 0**, valid JSON on stdout, no crash. The `/seo`
|
||||||
|
FULL audit continues on the anonymous PageSpeed API (lab data) instead of
|
||||||
|
CrUX field data, and the report surfaces the fix as a user action:
|
||||||
|
`make seo-connect`. `doctor.sh` also flags both non-fatally as `WARN`: a
|
||||||
|
missing `CRUX_API_KEY` warns on its own, while no connected Google account
|
||||||
|
is the one that names `make seo-connect`.
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make test
|
||||||
|
# or, to run only this engine's suite:
|
||||||
|
bash lib/seo-data/seo-data.test.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
The suite is network-free: `google_seo.py` reads fixtures from
|
||||||
|
`lib/seo-data/fixtures/` (`crux_mobile.json`, `gsc_queries.json`,
|
||||||
|
`gsc_inspect.json`) whenever `SEO_DATA_MOCK_DIR` is set, instead of calling
|
||||||
|
Google's APIs. Degradation paths run with real env vars unset (`env -u
|
||||||
|
CRUX_API_KEY`, `env -u SEO_DATA_MOCK_DIR`) to exercise the no-key/no-creds
|
||||||
|
branches deterministically. Every `fetch.sh` invocation in the tests also
|
||||||
|
sets `SEO_DATA_ENV_FILE=/dev/null` so a machine with a live
|
||||||
|
`~/.claude/.env` never lets real credentials leak into a test run.
|
||||||
@@ -0,0 +1,57 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""One-time OAuth consent + GSC property discovery + persist. Third-party imports
|
||||||
|
are lazy so `persist` is testable stdlib-only."""
|
||||||
|
import argparse, os, sys
|
||||||
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
import tokenstore
|
||||||
|
|
||||||
|
SCOPES = ["https://www.googleapis.com/auth/webmasters.readonly"]
|
||||||
|
|
||||||
|
def run_consent(client_id, client_secret):
|
||||||
|
from google_auth_oauthlib.flow import InstalledAppFlow # lazy
|
||||||
|
cfg = {"installed": {"client_id": client_id, "client_secret": client_secret,
|
||||||
|
"auth_uri": "https://accounts.google.com/o/oauth2/auth",
|
||||||
|
"token_uri": "https://oauth2.googleapis.com/token",
|
||||||
|
"redirect_uris": ["http://localhost"]}}
|
||||||
|
flow = InstalledAppFlow.from_client_config(cfg, scopes=SCOPES)
|
||||||
|
creds = flow.run_local_server(port=0) # opens browser, one-time consent
|
||||||
|
if not creds.refresh_token:
|
||||||
|
raise SystemExit("No refresh token returned. Revoke prior grant and retry.")
|
||||||
|
return creds.refresh_token
|
||||||
|
|
||||||
|
def discover_properties(refresh_token, client_id, client_secret):
|
||||||
|
from google.oauth2.credentials import Credentials
|
||||||
|
from google.auth.transport.requests import AuthorizedSession, Request
|
||||||
|
creds = Credentials(None, refresh_token=refresh_token, client_id=client_id,
|
||||||
|
client_secret=client_secret,
|
||||||
|
token_uri="https://oauth2.googleapis.com/token", scopes=SCOPES)
|
||||||
|
creds.refresh(Request())
|
||||||
|
r = AuthorizedSession(creds).get(
|
||||||
|
"https://searchconsole.googleapis.com/webmasters/v3/sites", timeout=30)
|
||||||
|
r.raise_for_status()
|
||||||
|
return [e["siteUrl"] for e in r.json().get("siteEntry", [])]
|
||||||
|
|
||||||
|
def persist(store_path, label, refresh_token, scopes, properties):
|
||||||
|
tokenstore.save_account(store_path, label, refresh_token, scopes, properties)
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
p.add_argument("--label", required=True)
|
||||||
|
p.add_argument("--store", default=os.path.expanduser("~/.claude/seo-data/tokens.json"))
|
||||||
|
args = p.parse_args()
|
||||||
|
cid = os.environ.get("GOOGLE_OAUTH_CLIENT_ID")
|
||||||
|
csec = os.environ.get("GOOGLE_OAUTH_CLIENT_SECRET")
|
||||||
|
if not (cid and csec):
|
||||||
|
raise SystemExit("Set GOOGLE_OAUTH_CLIENT_ID/SECRET in ~/.claude/.env first.")
|
||||||
|
existing = {a["label"] for a in tokenstore.list_accounts(args.store)}
|
||||||
|
if args.label in existing:
|
||||||
|
ans = input("Label '%s' exists. Overwrite? [y/N] " % args.label).strip().lower()
|
||||||
|
if ans != "y":
|
||||||
|
raise SystemExit("Aborted.")
|
||||||
|
rt = run_consent(cid, csec)
|
||||||
|
props = discover_properties(rt, cid, csec)
|
||||||
|
persist(args.store, args.label, rt, SCOPES, props)
|
||||||
|
print("Connected '%s'. Properties: %s" % (args.label, ", ".join(props) or "(none)"))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Stable entrypoint for the seo-data engine. JSON on stdout; exit 0 on ok/degrade,
|
||||||
|
# exit 2 on bad usage. Never prints secrets.
|
||||||
|
set -uo pipefail
|
||||||
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
ENV_FILE="${SEO_DATA_ENV_FILE:-${HOME}/.claude/.env}" # canonical; tests override to /dev/null
|
||||||
|
STORE="${SEO_DATA_STORE:-${HOME}/.claude/seo-data/tokens.json}"
|
||||||
|
VENV_PY="${HOME}/.claude/.venv-seo-data/bin/python3"
|
||||||
|
|
||||||
|
# Library stderr must never leak a secret into agent context — suppress it
|
||||||
|
# globally unless explicitly debugging (SEO_DATA_DEBUG=1 restores it).
|
||||||
|
[ -n "${SEO_DATA_DEBUG:-}" ] || exec 2>/dev/null
|
||||||
|
|
||||||
|
# Load secrets quietly (sourced, never echoed).
|
||||||
|
if [ -f "$ENV_FILE" ]; then
|
||||||
|
set -a; # shellcheck source=/dev/null
|
||||||
|
. "$ENV_FILE"; set +a
|
||||||
|
fi
|
||||||
|
# Prefer the isolated venv (has google-auth); fall back to system python3 for
|
||||||
|
# stdlib-only paths (accounts / mock / degrade).
|
||||||
|
PY="python3"; [ -x "$VENV_PY" ] && PY="$VENV_PY"
|
||||||
|
|
||||||
|
cmd="${1:-}"; shift || true
|
||||||
|
case "$cmd" in
|
||||||
|
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
||||||
|
crux|queries|inspect)
|
||||||
|
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
||||||
|
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect} [flags]"}'
|
||||||
|
exit 2 ;;
|
||||||
|
esac
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
{"record":{"key":{"formFactor":"PHONE"},"metrics":{
|
||||||
|
"largest_contentful_paint":{"percentiles":{"p75":2100}},
|
||||||
|
"interaction_to_next_paint":{"percentiles":{"p75":180}},
|
||||||
|
"cumulative_layout_shift":{"percentiles":{"p75":"0.08"}}}}}
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
{"inspectionResult":{"indexStatusResult":{
|
||||||
|
"verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}}
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
{"rows":[
|
||||||
|
{"keys":["plombier paris"],"clicks":40,"impressions":900,"ctr":0.044,"position":6.3},
|
||||||
|
{"keys":["urgence fuite"],"clicks":5,"impressions":1200,"ctr":0.004,"position":8.9}]}
|
||||||
@@ -0,0 +1,173 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""CrUX + GSC fetch → normalized JSON. Third-party imports are LAZY so mock and
|
||||||
|
degraded paths run stdlib-only (no venv, no network)."""
|
||||||
|
import argparse, json, os, sys
|
||||||
|
|
||||||
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
|
||||||
|
def _mock(name):
|
||||||
|
d = os.environ.get("SEO_DATA_MOCK_DIR")
|
||||||
|
if not d:
|
||||||
|
return None
|
||||||
|
path = os.path.join(d, name)
|
||||||
|
if not os.path.exists(path):
|
||||||
|
return None
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
|
||||||
|
def _norm_crux(raw):
|
||||||
|
m = raw["record"]["metrics"]
|
||||||
|
def p75(metric):
|
||||||
|
return m.get(metric, {}).get("percentiles", {}).get("p75")
|
||||||
|
out = {"status": "ok", "source": "crux"}
|
||||||
|
lcp = p75("largest_contentful_paint")
|
||||||
|
inp = p75("interaction_to_next_paint")
|
||||||
|
cls = p75("cumulative_layout_shift")
|
||||||
|
# Low-traffic origins often miss a metric (INP notably) — omit, don't crash.
|
||||||
|
if lcp is not None:
|
||||||
|
out["lcp_p75_ms"] = int(lcp)
|
||||||
|
if inp is not None:
|
||||||
|
out["inp_p75_ms"] = int(inp)
|
||||||
|
if cls is not None:
|
||||||
|
out["cls_p75"] = float(cls)
|
||||||
|
if len(out) == 2: # no metric at all
|
||||||
|
return {"status": "degraded", "reason": "no_field_data"}
|
||||||
|
return out
|
||||||
|
|
||||||
|
def _crux_query(key, body):
|
||||||
|
import requests # lazy
|
||||||
|
return requests.post(
|
||||||
|
"https://chromeuxreport.googleapis.com/v1/records:queryRecord?key=" + key,
|
||||||
|
json=body, timeout=20)
|
||||||
|
|
||||||
|
def _origin(url):
|
||||||
|
from urllib.parse import urlparse # stdlib
|
||||||
|
p = urlparse(url)
|
||||||
|
return "%s://%s" % (p.scheme, p.netloc) # strip path — CrUX origin = scheme+host only
|
||||||
|
|
||||||
|
def crux(url, strategy="mobile"):
|
||||||
|
raw = _mock("crux_%s.json" % strategy)
|
||||||
|
if raw is None:
|
||||||
|
key = os.environ.get("CRUX_API_KEY")
|
||||||
|
if not key:
|
||||||
|
return {"status": "degraded", "reason": "no_crux_key"}
|
||||||
|
ff = "PHONE" if strategy == "mobile" else "DESKTOP"
|
||||||
|
r = _crux_query(key, {"url": url, "formFactor": ff})
|
||||||
|
if r.status_code == 404: # no page-level data → try origin-level
|
||||||
|
r = _crux_query(key, {"origin": _origin(url), "formFactor": ff})
|
||||||
|
if r.status_code == 404:
|
||||||
|
return {"status": "degraded", "reason": "no_field_data"}
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
return _norm_crux(raw)
|
||||||
|
|
||||||
|
def _gsc_session(store_path, account):
|
||||||
|
"""Return an authorized requests.Session or a degrade dict. Lazy imports."""
|
||||||
|
rt = None
|
||||||
|
if store_path and account:
|
||||||
|
import tokenstore # local module, stdlib
|
||||||
|
rt = tokenstore.get_refresh_token(store_path, account)
|
||||||
|
cid = os.environ.get("GOOGLE_OAUTH_CLIENT_ID")
|
||||||
|
csec = os.environ.get("GOOGLE_OAUTH_CLIENT_SECRET")
|
||||||
|
if not (rt and cid and csec):
|
||||||
|
return {"status": "degraded", "reason": "no_credentials"}
|
||||||
|
from google.oauth2.credentials import Credentials # lazy
|
||||||
|
from google.auth.transport.requests import AuthorizedSession, Request
|
||||||
|
creds = Credentials(None, refresh_token=rt, client_id=cid, client_secret=csec,
|
||||||
|
token_uri="https://oauth2.googleapis.com/token",
|
||||||
|
scopes=["https://www.googleapis.com/auth/webmasters.readonly"])
|
||||||
|
try:
|
||||||
|
creds.refresh(Request())
|
||||||
|
except Exception as e:
|
||||||
|
# Only a real RefreshError means re-consent; a network blip must NOT
|
||||||
|
# send the user back through OAuth.
|
||||||
|
from google.auth.exceptions import RefreshError # lazy
|
||||||
|
reason = "token_revoked" if isinstance(e, RefreshError) else "network_error"
|
||||||
|
return {"status": "degraded", "reason": reason}
|
||||||
|
return AuthorizedSession(creds)
|
||||||
|
|
||||||
|
def _norm_queries(raw, dim):
|
||||||
|
return {"status": "ok", "source": "gsc", "dimension": dim, "rows": [
|
||||||
|
{"key": r["keys"][0], "clicks": r.get("clicks", 0),
|
||||||
|
"impressions": r.get("impressions", 0), "ctr": r.get("ctr", 0),
|
||||||
|
"position": r.get("position")}
|
||||||
|
for r in raw.get("rows", [])]}
|
||||||
|
|
||||||
|
def queries(store_path, account, property, days=90, dim="query"):
|
||||||
|
raw = _mock("gsc_queries.json")
|
||||||
|
if raw is None:
|
||||||
|
sess = _gsc_session(store_path, account)
|
||||||
|
if isinstance(sess, dict):
|
||||||
|
return sess
|
||||||
|
import datetime as _dt
|
||||||
|
end = _dt.date.today(); start = end - _dt.timedelta(days=days)
|
||||||
|
import urllib.parse
|
||||||
|
url = ("https://searchconsole.googleapis.com/webmasters/v3/sites/"
|
||||||
|
+ urllib.parse.quote(property, safe="") + "/searchAnalytics/query")
|
||||||
|
r = sess.post(url, json={"startDate": start.isoformat(), "endDate": end.isoformat(),
|
||||||
|
"dimensions": [dim], "rowLimit": 100}, timeout=30)
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
return _norm_queries(raw, dim)
|
||||||
|
|
||||||
|
def inspect(store_path, account, property, url):
|
||||||
|
raw = _mock("gsc_inspect.json")
|
||||||
|
if raw is None:
|
||||||
|
sess = _gsc_session(store_path, account)
|
||||||
|
if isinstance(sess, dict):
|
||||||
|
return sess
|
||||||
|
r = sess.post("https://searchconsole.googleapis.com/v1/urlInspection/index:inspect",
|
||||||
|
json={"inspectionUrl": url, "siteUrl": property}, timeout=30)
|
||||||
|
if r.status_code == 429:
|
||||||
|
return {"status": "degraded", "reason": "rate_limited"}
|
||||||
|
r.raise_for_status()
|
||||||
|
raw = r.json()
|
||||||
|
isr = raw["inspectionResult"]["indexStatusResult"]
|
||||||
|
return {"status": "ok", "source": "gsc",
|
||||||
|
"indexed": isr.get("verdict") == "PASS",
|
||||||
|
"coverage": isr.get("coverageState"),
|
||||||
|
"last_crawl": isr.get("lastCrawlTime")}
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
try:
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
sub = p.add_subparsers(dest="cmd", required=True)
|
||||||
|
pc = sub.add_parser("crux")
|
||||||
|
pc.add_argument("--url", required=True)
|
||||||
|
pc.add_argument("--strategy", default="mobile", choices=["mobile", "desktop"])
|
||||||
|
pc.add_argument("--store", default=None) # accepted+ignored: uniform fetch.sh dispatch
|
||||||
|
pq = sub.add_parser("queries")
|
||||||
|
pq.add_argument("--store", required=True)
|
||||||
|
pq.add_argument("--account", required=True)
|
||||||
|
pq.add_argument("--property", required=True)
|
||||||
|
pq.add_argument("--days", type=int, default=90)
|
||||||
|
pq.add_argument("--dim", default="query")
|
||||||
|
pi = sub.add_parser("inspect")
|
||||||
|
pi.add_argument("--store", required=True)
|
||||||
|
pi.add_argument("--account", required=True)
|
||||||
|
pi.add_argument("--property", required=True)
|
||||||
|
pi.add_argument("--url", required=True)
|
||||||
|
args = p.parse_args()
|
||||||
|
if args.cmd == "crux":
|
||||||
|
print(json.dumps(crux(args.url, args.strategy), indent=2))
|
||||||
|
elif args.cmd == "queries":
|
||||||
|
print(json.dumps(queries(args.store, args.account, args.property,
|
||||||
|
args.days, args.dim), indent=2))
|
||||||
|
elif args.cmd == "inspect":
|
||||||
|
print(json.dumps(inspect(args.store, args.account, args.property,
|
||||||
|
args.url), indent=2))
|
||||||
|
except SystemExit as e: # argparse usage error
|
||||||
|
if e.code not in (0, None):
|
||||||
|
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||||
|
raise # preserve argparse's exit code
|
||||||
|
except Exception:
|
||||||
|
# Fail-open data contract: ANY unexpected error (HTTP 403/5xx, DNS,
|
||||||
|
# timeout) degrades with exit 0 — never a traceback, never empty stdout.
|
||||||
|
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
google-auth==2.40.0
|
||||||
|
google-auth-oauthlib==1.2.2
|
||||||
|
requests==2.32.4
|
||||||
@@ -0,0 +1,133 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Deterministic tests for the seo-data engine (no network, no venv).
|
||||||
|
set -u
|
||||||
|
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||||
|
SD="$REPO/lib/seo-data"
|
||||||
|
PASS=0; FAIL=0
|
||||||
|
ok() { echo " PASS $1"; PASS=$((PASS+1)); }
|
||||||
|
no() { echo " FAIL $1 — $2"; FAIL=$((FAIL+1)); }
|
||||||
|
# assert stdout of a command contains / omits a fixed string
|
||||||
|
has() { if printf '%s' "$2" | grep -qF -- "$3"; then ok "$1"; else no "$1" "missing: $3"; fi; }
|
||||||
|
hasnt(){ if printf '%s' "$2" | grep -qF -- "$3"; then no "$1" "forbidden: $3"; else ok "$1"; fi; }
|
||||||
|
|
||||||
|
echo "── tokenstore ──"
|
||||||
|
TMP="$(mktemp -d)"; STORE="$TMP/tokens.json"
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$STORE" --label client-a \
|
||||||
|
--refresh-token RT_AAA --scopes https://www.googleapis.com/auth/webmasters.readonly \
|
||||||
|
--properties sc-domain:a.com,https://www.a.com/ >/dev/null
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$STORE" --label client-b \
|
||||||
|
--refresh-token RT_BBB --scopes https://www.googleapis.com/auth/webmasters.readonly \
|
||||||
|
--properties sc-domain:b.com >/dev/null
|
||||||
|
LIST="$(python3 "$SD/tokenstore.py" list --file "$STORE")"
|
||||||
|
has "list shows client-a" "$LIST" '"client-a"'
|
||||||
|
has "list shows client-b" "$LIST" '"client-b"'
|
||||||
|
has "list shows a property" "$LIST" 'sc-domain:a.com'
|
||||||
|
hasnt "list redacts refresh tokens" "$LIST" 'RT_AAA'
|
||||||
|
PERM="$(stat -c '%a' "$STORE")"
|
||||||
|
[ "$PERM" = "600" ] && ok "store file is 0600" || no "store file 0600" "got $PERM"
|
||||||
|
DPERM="$(stat -c '%a' "$(dirname "$STORE")")"
|
||||||
|
[ "$DPERM" = "700" ] && ok "store dir is 0700" || no "store dir 0700" "got $DPERM"
|
||||||
|
rm -rf "$TMP"
|
||||||
|
|
||||||
|
echo "── crux (mock) ──"
|
||||||
|
CRUX_OK="$(SEO_DATA_MOCK_DIR="$REPO/lib/seo-data/fixtures" \
|
||||||
|
python3 "$SD/google_seo.py" crux --url https://ex.com --strategy mobile)"
|
||||||
|
has "crux status ok" "$CRUX_OK" '"status": "ok"'
|
||||||
|
has "crux lcp p75 mapped" "$CRUX_OK" '"lcp_p75_ms": 2100'
|
||||||
|
has "crux inp p75 mapped" "$CRUX_OK" '"inp_p75_ms": 180'
|
||||||
|
has "crux cls p75 mapped" "$CRUX_OK" '"cls_p75": 0.08'
|
||||||
|
CRUX_DEG="$(env -u CRUX_API_KEY -u SEO_DATA_MOCK_DIR \
|
||||||
|
python3 "$SD/google_seo.py" crux --url https://ex.com)"
|
||||||
|
has "crux degrades w/o key" "$CRUX_DEG" '"status": "degraded"'
|
||||||
|
has "crux degrade reason" "$CRUX_DEG" 'no_crux_key'
|
||||||
|
ORIG="$(python3 -c "import sys; sys.path.insert(0,'$SD'); import google_seo; print(google_seo._origin('https://example.com/blog/post'))")"
|
||||||
|
has "origin strips to host" "$ORIG" 'https://example.com'
|
||||||
|
hasnt "origin drops the path" "$ORIG" 'blog'
|
||||||
|
|
||||||
|
echo "── gsc (mock) ──"
|
||||||
|
MOCK="$REPO/lib/seo-data/fixtures"
|
||||||
|
TMP2="$(mktemp -d)"; S2="$TMP2/tokens.json"
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$S2" --label client-a --refresh-token RT \
|
||||||
|
--scopes https://www.googleapis.com/auth/webmasters.readonly --properties sc-domain:ex.com >/dev/null
|
||||||
|
Q="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" queries \
|
||||||
|
--store "$S2" --account client-a --property sc-domain:ex.com --days 90)"
|
||||||
|
has "queries ok" "$Q" '"status": "ok"'
|
||||||
|
has "queries row key" "$Q" 'plombier paris'
|
||||||
|
has "queries position field" "$Q" '"position": 6.3'
|
||||||
|
I="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" inspect \
|
||||||
|
--store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)"
|
||||||
|
has "inspect indexed true" "$I" '"indexed": true'
|
||||||
|
DEG="$(env -u SEO_DATA_MOCK_DIR python3 "$SD/google_seo.py" queries \
|
||||||
|
--store "$TMP2/none.json" --account nobody --property sc-domain:ex.com)"
|
||||||
|
has "gsc degrades w/o creds" "$DEG" '"status": "degraded"'
|
||||||
|
has "gsc degrade reason" "$DEG" 'no_credentials'
|
||||||
|
rm -rf "$TMP2"
|
||||||
|
|
||||||
|
echo "── fetch.sh ──"
|
||||||
|
FETCH="$SD/fetch.sh"
|
||||||
|
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
|
||||||
|
# on a machine with a live CRUX_API_KEY the degrade tests would hit the network.
|
||||||
|
NOENV=/dev/null
|
||||||
|
ACC="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_STORE=/nonexistent/tokens.json bash "$FETCH" accounts)"
|
||||||
|
has "accounts empty is ok json" "$ACC" '"accounts": []'
|
||||||
|
CR="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_MOCK_DIR="$MOCK" bash "$FETCH" crux --url https://ex.com)"
|
||||||
|
has "fetch crux ok" "$CR" '"status": "ok"'
|
||||||
|
SEO_DATA_ENV_FILE=$NOENV bash "$FETCH" bogus-subcmd >/dev/null 2>&1; RC=$?
|
||||||
|
[ "$RC" = "2" ] && ok "bad subcmd exit 2" || no "bad subcmd exit 2" "got $RC"
|
||||||
|
DG="$(SEO_DATA_ENV_FILE=$NOENV env -u SEO_DATA_MOCK_DIR -u CRUX_API_KEY bash "$FETCH" crux --url https://ex.com)"; RC=$?
|
||||||
|
has "degrade json" "$DG" '"status": "degraded"'
|
||||||
|
[ "$RC" = "0" ] && ok "degrade exit 0" || no "degrade exit 0" "got $RC"
|
||||||
|
# redaction through the real fetch.sh dispatch layer
|
||||||
|
TMP4="$(mktemp -d)"; RSTORE="$TMP4/rt.json"
|
||||||
|
python3 "$SD/tokenstore.py" set --file "$RSTORE" --label leaky --refresh-token RT_SECRET_XYZ \
|
||||||
|
--scopes https://www.googleapis.com/auth/webmasters.readonly --properties sc-domain:z.com >/dev/null
|
||||||
|
ACCJSON="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_STORE="$RSTORE" bash "$FETCH" accounts)"
|
||||||
|
has "accounts lists label" "$ACCJSON" 'leaky'
|
||||||
|
hasnt "accounts hides token" "$ACCJSON" 'RT_SECRET_XYZ'
|
||||||
|
rm -rf "$TMP4"
|
||||||
|
# corrupted store must degrade with JSON + exit 0 (Fix 1)
|
||||||
|
TMP5="$(mktemp -d)"; CSTORE="$TMP5/corrupt.json"; printf 'not json {{' > "$CSTORE"
|
||||||
|
CJ="$(SEO_DATA_ENV_FILE=$NOENV SEO_DATA_STORE="$CSTORE" bash "$FETCH" accounts)"; CRC=$?
|
||||||
|
has "corrupt store degrades" "$CJ" '"status"'
|
||||||
|
[ "$CRC" = "0" ] && ok "corrupt store exit 0" || no "corrupt store exit 0" "got $CRC"
|
||||||
|
rm -rf "$TMP5"
|
||||||
|
# bad usage (known subcmd, missing flag) must still emit JSON + exit 2 (Fix 2)
|
||||||
|
BU="$(SEO_DATA_ENV_FILE=$NOENV bash "$FETCH" crux)"; BURC=$?
|
||||||
|
has "bad usage emits json" "$BU" '"status"'
|
||||||
|
[ "$BURC" = "2" ] && ok "bad usage exit 2" || no "bad usage exit 2" "got $BURC"
|
||||||
|
|
||||||
|
echo "── connect (persist, offline) ──"
|
||||||
|
TMP3="$(mktemp -d)"; S3="$TMP3/tokens.json"
|
||||||
|
python3 -c "import sys; sys.path.insert(0,'$SD'); import connect; \
|
||||||
|
connect.persist('$S3','client-x','RT_X',['https://www.googleapis.com/auth/webmasters.readonly'],['sc-domain:x.com'])"
|
||||||
|
L3="$(python3 "$SD/tokenstore.py" list --file "$S3")"
|
||||||
|
has "connect.persist wrote label" "$L3" '"client-x"'
|
||||||
|
has "connect.persist wrote prop" "$L3" 'sc-domain:x.com'
|
||||||
|
hasnt "connect.persist redacts" "$L3" 'RT_X'
|
||||||
|
rm -rf "$TMP3"
|
||||||
|
|
||||||
|
echo "── wiring locks ──"
|
||||||
|
tf() { if grep -qF -- "$3" "$2" 2>/dev/null; then ok "$1"; else no "$1" "missing: $3"; fi; }
|
||||||
|
tf "env.example client id" "$REPO/.env.example" "GOOGLE_OAUTH_CLIENT_ID="
|
||||||
|
tf "env.example crux key" "$REPO/.env.example" "CRUX_API_KEY="
|
||||||
|
tf "makefile seo-connect" "$REPO/Makefile" "seo-connect:"
|
||||||
|
tf "makefile discovers test" "$REPO/Makefile" "lib/seo-data/*.test.sh"
|
||||||
|
tf "install prompts connect" "$REPO/install.sh" "make seo-connect"
|
||||||
|
tf "doctor checks seo-data" "$REPO/doctor.sh" "seo-data"
|
||||||
|
tf "gitleaks allowlist store" "$REPO/.gitleaks.toml" "seo-data/tokens"
|
||||||
|
tf "gitignore venv" "$REPO/.gitignore" ".venv-seo-data"
|
||||||
|
|
||||||
|
echo "── integration locks ──"
|
||||||
|
tf "skill step0 account select" "$REPO/skills/seo/SKILL.md" "COMPTE GOOGLE"
|
||||||
|
tf "analyzer calls fetch crux" "$REPO/agents/seo-analyzer.md" "fetch.sh crux"
|
||||||
|
tf "analyzer calls fetch queries" "$REPO/agents/seo-analyzer.md" "fetch.sh queries"
|
||||||
|
tf "analyzer gsc subsection" "$REPO/agents/seo-analyzer.md" "Performance GSC"
|
||||||
|
tf "catalog gsc oauth entry" "$REPO/agents/resources/automation-catalog.md" "make seo-connect"
|
||||||
|
|
||||||
|
echo "── readme lock ──"
|
||||||
|
tf "readme documents fetch.sh" "$REPO/lib/seo-data/README.md" "fetch.sh"
|
||||||
|
tf "readme documents seo-connect" "$REPO/lib/seo-data/README.md" "make seo-connect"
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "seo-data engine: $PASS pass, $FAIL fail"
|
||||||
|
[ "$FAIL" -eq 0 ]
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Label-keyed OAuth refresh-token store. Atomic writes under an fcntl lock.
|
||||||
|
No third-party deps — must run without the venv (used by the offline test path)."""
|
||||||
|
import argparse, fcntl, json, os, tempfile
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
def load(path):
|
||||||
|
if not os.path.exists(path):
|
||||||
|
return {"version": 1, "accounts": {}}
|
||||||
|
with open(path, "r", encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
|
||||||
|
def list_accounts(path):
|
||||||
|
data = load(path)
|
||||||
|
return [
|
||||||
|
{"label": lbl, "properties": a.get("properties", []),
|
||||||
|
"granted_at": a.get("granted_at")}
|
||||||
|
for lbl, a in data.get("accounts", {}).items()
|
||||||
|
] # refresh_token intentionally omitted (redaction)
|
||||||
|
|
||||||
|
def get_refresh_token(path, label):
|
||||||
|
return load(path).get("accounts", {}).get(label, {}).get("refresh_token")
|
||||||
|
|
||||||
|
def save_account(path, label, refresh_token, scopes, properties):
|
||||||
|
dirpath = os.path.dirname(path) or "."
|
||||||
|
os.makedirs(dirpath, mode=0o700, exist_ok=True)
|
||||||
|
os.chmod(dirpath, 0o700) # re-assert invariant (makedirs no-ops if dir exists)
|
||||||
|
lock_path = path + ".lock"
|
||||||
|
with open(lock_path, "w") as lock:
|
||||||
|
os.chmod(lock_path, 0o600) # defense-in-depth (empty flock handle, never holds token)
|
||||||
|
fcntl.flock(lock, fcntl.LOCK_EX) # serialize concurrent connects
|
||||||
|
data = load(path)
|
||||||
|
data.setdefault("version", 1)
|
||||||
|
data.setdefault("accounts", {})
|
||||||
|
data["accounts"][label] = {
|
||||||
|
"refresh_token": refresh_token,
|
||||||
|
"scopes": scopes,
|
||||||
|
"granted_at": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"properties": properties,
|
||||||
|
}
|
||||||
|
fd, tmp = tempfile.mkstemp(dir=dirpath, suffix=".tmp")
|
||||||
|
try:
|
||||||
|
with os.fdopen(fd, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(data, f, indent=2)
|
||||||
|
f.flush(); os.fsync(f.fileno())
|
||||||
|
os.chmod(tmp, 0o600)
|
||||||
|
os.replace(tmp, path) # atomic
|
||||||
|
finally:
|
||||||
|
if os.path.exists(tmp):
|
||||||
|
os.unlink(tmp)
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
sub = p.add_subparsers(dest="cmd", required=True)
|
||||||
|
pl = sub.add_parser("list"); pl.add_argument("--file", required=True)
|
||||||
|
ps = sub.add_parser("set")
|
||||||
|
for flag in ("--file", "--label", "--refresh-token"):
|
||||||
|
ps.add_argument(flag, required=True)
|
||||||
|
ps.add_argument("--scopes", default="")
|
||||||
|
ps.add_argument("--properties", default="")
|
||||||
|
try:
|
||||||
|
args = p.parse_args()
|
||||||
|
if args.cmd == "list":
|
||||||
|
print(json.dumps({"status": "ok", "accounts": list_accounts(args.file)}))
|
||||||
|
else:
|
||||||
|
save_account(args.file, args.label, getattr(args, "refresh_token"),
|
||||||
|
[s for s in args.scopes.split(",") if s],
|
||||||
|
[x for x in args.properties.split(",") if x])
|
||||||
|
print(json.dumps({"status": "ok"}))
|
||||||
|
except SystemExit as e: # argparse usage error
|
||||||
|
if e.code not in (0, None):
|
||||||
|
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||||
|
raise # preserve argparse's exit code
|
||||||
|
except Exception: # e.g. corrupted store JSON
|
||||||
|
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
@@ -63,6 +63,43 @@ If `$ARGUMENTS` contains `local`/`code-only`/`quick`/`rapide` → default LOCAL.
|
|||||||
If `$ARGUMENTS` contains `full`/`complet`/`externe`/`live` → default FULL.
|
If `$ARGUMENTS` contains `full`/`complet`/`externe`/`live` → default FULL.
|
||||||
If `$ARGUMENTS` contains a production URL → suggest FULL.
|
If `$ARGUMENTS` contains a production URL → suggest FULL.
|
||||||
|
|
||||||
|
### Compte Google (FULL only)
|
||||||
|
|
||||||
|
**Skip if LOCAL** — jump straight to Business context.
|
||||||
|
|
||||||
|
For FULL depth, offer to attach a connected Google account so the
|
||||||
|
seo-analyzer can pull real GSC/CrUX data instead of anonymous PageSpeed
|
||||||
|
only. List connected accounts (**tilde path mandatory** — this skill
|
||||||
|
runs from the audited project's directory, not the claude-config repo):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh accounts
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
COMPTE GOOGLE pour cet audit FULL :
|
||||||
|
|
||||||
|
1. <label> — <property 1>, <property 2>, ...
|
||||||
|
2. <label> — <property>
|
||||||
|
...
|
||||||
|
[connecter un nouveau compte] — lancer `make seo-connect` (depuis le
|
||||||
|
repo claude-config, une fois par compte), puis relancer /seo
|
||||||
|
[Ignorer] — continuer sans GSC/CrUX (PageSpeed anonyme uniquement,
|
||||||
|
dégradation normale — cf. SEO.md §11)
|
||||||
|
|
||||||
|
Quel compte / quelle property ? (numéro, ou "ignorer")
|
||||||
|
```
|
||||||
|
|
||||||
|
If `fetch.sh accounts` returns an empty list (`"accounts": []`), skip
|
||||||
|
the numbered list and show only `[connecter un nouveau compte]` /
|
||||||
|
`[Ignorer]`.
|
||||||
|
|
||||||
|
Record the choice in the shared context block:
|
||||||
|
```
|
||||||
|
GSC ACCOUNT: <label> | none
|
||||||
|
GSC PROPERTY: <property> | none
|
||||||
|
```
|
||||||
|
|
||||||
### Business context (one grouped block)
|
### Business context (one grouped block)
|
||||||
|
|
||||||
**Both depths:**
|
**Both depths:**
|
||||||
@@ -172,6 +209,8 @@ BUSINESS CONTEXT:
|
|||||||
Known citations: ...
|
Known citations: ...
|
||||||
Known competitors: ...
|
Known competitors: ...
|
||||||
Time budget: ...
|
Time budget: ...
|
||||||
|
GSC account: <label> | none (FULL only)
|
||||||
|
GSC property: <property> | none (FULL only)
|
||||||
|
|
||||||
You are the classical-SEO half of a parallel SEO+GEO audit. Do NOT
|
You are the classical-SEO half of a parallel SEO+GEO audit. Do NOT
|
||||||
audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas,
|
audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas,
|
||||||
|
|||||||
Reference in New Issue
Block a user