From a6d423b940740055f59de4e226629a70489576f7 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:51:30 +0200 Subject: [PATCH] =?UTF-8?q?feat(seo-data):=20W1=20=E2=80=94=20surface=20ri?= =?UTF-8?q?ch=5Fresults,=20the=20data=20inspect=20already=20threw=20away?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit google_seo.py:129 read only indexStatusResult out of the URL Inspection response and discarded the rest. richResultsResult was already on the wire: same call, same OAuth scope (webmasters.readonly), same quota. Google's own structured-data verdict on the live indexed URL was being downloaded and binned. Plan correction: the TODO said "richresults verb". Wrong — a new verb means a second POST to the same endpoint for a payload already received, on a per-site quota, and nobody wants rich results without index status. Extended inspect() instead; fetch.sh unchanged, no new verb, no new scope. Design driven by the published schema, not by guesswork — two details I would have got wrong: - richResultsResult is OMITTED when Google detects none ("absent if none found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a caller cannot tell an absent key from a check that never ran. ABSENT means "none detected", never "invalid". The KeyError path is the real risk here, so it has its own fixture dir (fixtures-norich/) and its own tests. - PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It never emits it now, and a test asserts the absence. issues[] deduped (the same issueMessage repeats across every affected item), errors/warnings count instances — scale from the counter, cause from the message. seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a verified property, so its reach is the STEP 9 COVERAGE ratio, not the site. Replacing a fake validator with a fake coverage promise would be no better. This is what beats claude-seo: their README's "dual validator (Rich Results Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of their .py finds zero calls. This is Google's verdict, via auth already held. Note: the new dedupe assertion trips SC2015 (A && B || C), same as the pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh shellcheck glob anyway. Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean. --- agents/seo-analyzer.md | 29 +++++++++++++ lib/seo-data/README.md | 17 +++++++- lib/seo-data/fixtures-norich/gsc_inspect.json | 2 + lib/seo-data/fixtures/gsc_inspect.json | 13 +++++- lib/seo-data/google_seo.py | 41 ++++++++++++++++++- lib/seo-data/seo-data.test.sh | 18 ++++++++ 6 files changed, 115 insertions(+), 5 deletions(-) create mode 100644 lib/seo-data/fixtures-norich/gsc_inspect.json diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 190a865..8c7070d 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -339,6 +339,35 @@ and 10 AND high impressions (candidates to push onto page 1 with a title/meta/content tweak). Report index coverage from `inspect`. All emitted into SEO.md §2 (technical) and §8 (quick wins). +**`inspect` also returns `rich_results` — Google's own structured-data +verdict on the live indexed URL.** It rides the same response (no extra +call, no extra quota). This is the only programmatic JSON-LD validation in +the system; everything else about schema is read by eye. + +``` +rich_results.verdict : PASS | FAIL | NEUTRAL | VERDICT_UNSPECIFIED | ABSENT +rich_results.types[] : {type, items, errors, warnings, issues[]} +``` + +- `FAIL` + a type carrying `errors > 0` → that type **cannot show as a rich + result**. Bundle item, cite the `issues[]` message verbatim — it is + Google's wording, not ours, and geo-analyzer owns the JSON-LD fix + (CROSS-AGENT NOTE). +- `warnings` → recommended fields missing. Report, do not gate on them. +- **`ABSENT` means Google detected no rich results on this URL** — the key + is omitted upstream when nothing is found. It is NOT an error and NOT + proof the markup is broken: a page with no structured data reads the same + as one whose markup Google never parsed. Say "none detected", never + "invalid". +- `ABSENT` while the repo clearly ships JSON-LD → real finding: the markup + is not reaching Google (SPA-rendered, blocked, or malformed). Cross-check + before claiming it. + +**Bound this honestly.** `index:inspect` is per-URL, quota'd, and works only +on a GSC-verified property. It validates the URLs you sampled — not the +site. Its reach is the STEP 9 COVERAGE ratio, and §14 must say so rather +than let one PASS imply site-wide valid markup. + If `status=degraded` → note it in §2 and emit the §11 user action "Connecter GSC: `make seo-connect`". diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index b730beb..c21b097 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -80,9 +80,24 @@ fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--d → {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"} fetch.sh inspect --account client-a --property … --url https://ex.com/page - → {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…"} + → {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…", + "rich_results":{"verdict":"PASS|FAIL|NEUTRAL|VERDICT_UNSPECIFIED|ABSENT", + "types":[{"type":"FAQ","items":2,"errors":2,"warnings":1, + "issues":["Missing field 'acceptedAnswer'"]}]}} → {"status":"degraded","reason":"…"} + rich_results rides the SAME URL-Inspection response — Google already sends + it, `inspect` used to discard it. No extra call, quota or OAuth scope. + It is the only programmatic structured-data validation in the system. + • verdict PARTIAL is never emitted — the API reserves it as unused. + • verdict ABSENT is SYNTHETIC (not a Google enum): the API omits + richResultsResult entirely when it detects no rich results. Surfaced + as a value rather than a missing key, because a caller cannot tell an + absent key apart from a check that never ran. ABSENT = "none + detected", never "invalid". + • errors/warnings count issue INSTANCES; issues[] is deduped — the same + issueMessage repeats across every affected item. + fetch.sh forget --label client-a → {"status":"ok","removed":true|false} # false = label wasn't in the store diff --git a/lib/seo-data/fixtures-norich/gsc_inspect.json b/lib/seo-data/fixtures-norich/gsc_inspect.json new file mode 100644 index 0000000..325cacd --- /dev/null +++ b/lib/seo-data/fixtures-norich/gsc_inspect.json @@ -0,0 +1,2 @@ +{"inspectionResult":{"indexStatusResult":{ + "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}} diff --git a/lib/seo-data/fixtures/gsc_inspect.json b/lib/seo-data/fixtures/gsc_inspect.json index 325cacd..bb5de1f 100644 --- a/lib/seo-data/fixtures/gsc_inspect.json +++ b/lib/seo-data/fixtures/gsc_inspect.json @@ -1,2 +1,11 @@ -{"inspectionResult":{"indexStatusResult":{ - "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}} +{"inspectionResult":{ + "indexStatusResult":{ + "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}, + "richResultsResult":{"verdict":"FAIL","detectedItems":[ + {"richResultType":"Breadcrumbs","items":[{"name":"Unnamed item","issues":[]}]}, + {"richResultType":"FAQ","items":[ + {"name":"Q1","issues":[ + {"issueMessage":"Missing field 'acceptedAnswer'","severity":"ERROR"}]}, + {"name":"Q2","issues":[ + {"issueMessage":"Missing field 'acceptedAnswer'","severity":"ERROR"}, + {"issueMessage":"Unspecified image","severity":"WARNING"}]}]}]}}} diff --git a/lib/seo-data/google_seo.py b/lib/seo-data/google_seo.py index d73277d..a544c6f 100644 --- a/lib/seo-data/google_seo.py +++ b/lib/seo-data/google_seo.py @@ -114,6 +114,39 @@ def queries(store_path, account, property, days=90, dim="query"): raw = r.json() return _norm_queries(raw, dim) +def _rollup_issues(items): + """Count issue instances by severity; dedupe messages (they repeat per item).""" + errors = warnings = 0 + msgs = [] + for item in items: + for iss in item.get("issues", []): + sev = iss.get("severity") + if sev == "ERROR": + errors += 1 + elif sev == "WARNING": + warnings += 1 + msg = iss.get("issueMessage") + if msg and msg not in msgs: + msgs.append(msg) + return errors, warnings, msgs + +def _norm_rich(ir): + """richResultsResult → verdict + per-type rollup. Google OMITS the key when + it detects no rich results, so absence is data, not an error: surfaced as the + synthetic verdict ABSENT (not a Google enum) rather than a missing key, which + a caller cannot tell apart from a check that never ran. PARTIAL is never + emitted — the API reserves it as unused.""" + rr = ir.get("richResultsResult") + if rr is None: + return {"verdict": "ABSENT", "types": []} + types = [] + for det in rr.get("detectedItems", []): + errors, warnings, msgs = _rollup_issues(det.get("items", [])) + types.append({"type": det.get("richResultType"), + "items": len(det.get("items", [])), + "errors": errors, "warnings": warnings, "issues": msgs}) + return {"verdict": rr.get("verdict"), "types": types} + def inspect(store_path, account, property, url): raw = _mock("gsc_inspect.json") if raw is None: @@ -126,11 +159,15 @@ def inspect(store_path, account, property, url): return {"status": "degraded", "reason": "rate_limited"} r.raise_for_status() raw = r.json() - isr = raw["inspectionResult"]["indexStatusResult"] + ir = raw["inspectionResult"] + isr = ir["indexStatusResult"] + # rich_results rides the SAME response — Google already sent it and this + # function used to discard it. No extra call, no extra quota, no new scope. return {"status": "ok", "source": "gsc", "indexed": isr.get("verdict") == "PASS", "coverage": isr.get("coverageState"), - "last_crawl": isr.get("lastCrawlTime")} + "last_crawl": isr.get("lastCrawlTime"), + "rich_results": _norm_rich(ir)} def _cli(): try: diff --git a/lib/seo-data/seo-data.test.sh b/lib/seo-data/seo-data.test.sh index 16e8a81..751a8e0 100644 --- a/lib/seo-data/seo-data.test.sh +++ b/lib/seo-data/seo-data.test.sh @@ -57,6 +57,24 @@ has "queries position field" "$Q" '"position": 6.3' I="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" inspect \ --store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)" has "inspect indexed true" "$I" '"indexed": true' +# rich_results rides the same URL-Inspection response (no extra call/quota) +has "rich verdict surfaced" "$I" '"verdict": "FAIL"' +has "rich type breadcrumbs" "$I" '"type": "Breadcrumbs"' +has "rich type faq" "$I" '"type": "FAQ"' +has "rich counts error severity" "$I" '"errors": 2' +has "rich counts warn severity" "$I" '"warnings": 1' +has "rich keeps issue message" "$I" "Missing field 'acceptedAnswer'" +# same issueMessage repeats across items — the rollup must collapse it to one +NMSG="$(printf '%s' "$I" | grep -cF "Missing field 'acceptedAnswer'")" +[ "$NMSG" = "1" ] && ok "rich dedupes issue messages" \ + || no "rich dedupes issue messages" "got $NMSG occurrences" +# Google OMITS richResultsResult when it detects none — absence is data, and +# must not KeyError nor vanish into a missing key +NR="$(SEO_DATA_MOCK_DIR="$SD/fixtures-norich" python3 "$SD/google_seo.py" inspect \ + --store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)" +has "no-rich → synthetic ABSENT" "$NR" '"verdict": "ABSENT"' +has "no-rich keeps index status" "$NR" '"indexed": true' +hasnt "no-rich emits no PARTIAL" "$NR" 'PARTIAL' DEG="$(env -u SEO_DATA_MOCK_DIR python3 "$SD/google_seo.py" queries \ --store "$TMP2/none.json" --account nobody --property sc-domain:ex.com)" has "gsc degrades w/o creds" "$DEG" '"status": "degraded"'