feat(seo-data): W1 — surface rich_results, the data inspect already threw away

google_seo.py:129 read only indexStatusResult out of the URL Inspection
response and discarded the rest. richResultsResult was already on the wire:
same call, same OAuth scope (webmasters.readonly), same quota. Google's own
structured-data verdict on the live indexed URL was being downloaded and
binned.

Plan correction: the TODO said "richresults verb". Wrong — a new verb means
a second POST to the same endpoint for a payload already received, on a
per-site quota, and nobody wants rich results without index status. Extended
inspect() instead; fetch.sh unchanged, no new verb, no new scope.

Design driven by the published schema, not by guesswork — two details I
would have got wrong:
- richResultsResult is OMITTED when Google detects none ("absent if none
  found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a
  caller cannot tell an absent key from a check that never ran. ABSENT means
  "none detected", never "invalid". The KeyError path is the real risk here,
  so it has its own fixture dir (fixtures-norich/) and its own tests.
- PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It
  never emits it now, and a test asserts the absence.

issues[] deduped (the same issueMessage repeats across every affected item),
errors/warnings count instances — scale from the counter, cause from the
message.

seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD
validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a
verified property, so its reach is the STEP 9 COVERAGE ratio, not the site.
Replacing a fake validator with a fake coverage promise would be no better.

This is what beats claude-seo: their README's "dual validator (Rich Results
Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of
their .py finds zero calls. This is Google's verdict, via auth already held.

Note: the new dedupe assertion trips SC2015 (A && B || C), same as the
pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for
house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh
shellcheck glob anyway.

Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end
and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-16 20:51:30 +02:00
parent fe93b7945b
commit a6d423b940
6 changed files with 115 additions and 5 deletions
+39 -2
View File
@@ -114,6 +114,39 @@ def queries(store_path, account, property, days=90, dim="query"):
raw = r.json()
return _norm_queries(raw, dim)
def _rollup_issues(items):
"""Count issue instances by severity; dedupe messages (they repeat per item)."""
errors = warnings = 0
msgs = []
for item in items:
for iss in item.get("issues", []):
sev = iss.get("severity")
if sev == "ERROR":
errors += 1
elif sev == "WARNING":
warnings += 1
msg = iss.get("issueMessage")
if msg and msg not in msgs:
msgs.append(msg)
return errors, warnings, msgs
def _norm_rich(ir):
"""richResultsResult → verdict + per-type rollup. Google OMITS the key when
it detects no rich results, so absence is data, not an error: surfaced as the
synthetic verdict ABSENT (not a Google enum) rather than a missing key, which
a caller cannot tell apart from a check that never ran. PARTIAL is never
emitted — the API reserves it as unused."""
rr = ir.get("richResultsResult")
if rr is None:
return {"verdict": "ABSENT", "types": []}
types = []
for det in rr.get("detectedItems", []):
errors, warnings, msgs = _rollup_issues(det.get("items", []))
types.append({"type": det.get("richResultType"),
"items": len(det.get("items", [])),
"errors": errors, "warnings": warnings, "issues": msgs})
return {"verdict": rr.get("verdict"), "types": types}
def inspect(store_path, account, property, url):
raw = _mock("gsc_inspect.json")
if raw is None:
@@ -126,11 +159,15 @@ def inspect(store_path, account, property, url):
return {"status": "degraded", "reason": "rate_limited"}
r.raise_for_status()
raw = r.json()
isr = raw["inspectionResult"]["indexStatusResult"]
ir = raw["inspectionResult"]
isr = ir["indexStatusResult"]
# rich_results rides the SAME response — Google already sent it and this
# function used to discard it. No extra call, no extra quota, no new scope.
return {"status": "ok", "source": "gsc",
"indexed": isr.get("verdict") == "PASS",
"coverage": isr.get("coverageState"),
"last_crawl": isr.get("lastCrawlTime")}
"last_crawl": isr.get("lastCrawlTime"),
"rich_results": _norm_rich(ir)}
def _cli():
try: