feat(seo-data): schema_gen verb — generate JSON-LD, not just audit it

Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.

fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.

Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).

geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.

Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 14:30:31 +02:00
parent 92301fe1c8
commit cfdd89e73b
5 changed files with 404 additions and 2 deletions
+33
View File
@@ -215,6 +215,39 @@ fetch.sh score --findings <path.json | ->
• Malformed input is an error, never a silently wrong number — unlike the
fetch verbs, a degrade here would mean bad input, not a network fact.
fetch.sh schema_gen <reservation|order|discussion|profile> [flags] [--script-tag]
→ {"status":"ok","source":"schema_gen","type":"<@type>","jsonld":{…}}
→ {"status":"error","reason":"bad_usage"} # a REQUIRED flag omitted
→ {"status":"degraded","reason":"…"} # a required flag given, empty
fetch.sh schema_gen reservation --provider "Marea NYC" \
--start 2026-06-04T19:30:00-04:00 --party-size 4
fetch.sh schema_gen order --merchant "Acme Pizza" --order-url https://acme.example/order
fetch.sh schema_gen discussion --headline "…" --author "Sara Park" \
--url https://forum.example.com/t/123 --date 2026-05-12T14:00:00Z
fetch.sh schema_gen profile --name "Daniel Agrici" --url https://agricidaniel.com/about \
--same-as https://github.com/AgriciDaniel --knows-about "SEO" "Schema markup"
Adapted from claude-seo's `schema_generate.py` (MIT) into this contract.
Our system only AUDITS existing markup elsewhere; this is the one verb
that GENERATES it — deterministic JSON-LD skeletons for the four v2
high-leverage Schema.org types, so geo-analyzer's G2 batch stops
hand-writing markup by hand. It only generates STRUCTURE: unknown field
VALUES are the caller's job, `[À COMPLÉTER]` for anything unconfirmed —
this verb never invents a sameAs, an email, or a business name.
• Stdlib only, no network, no auth — runs even without the venv.
• `--script-tag` wraps the cleaned jsonld in
`<script type="application/ld+json">…</script>` under a `script` key,
still inside the `ok` envelope. It must be given AFTER the type
(`schema_gen reservation … --script-tag`, not before) — argparse
subcommand flags only parse after their subcommand.
• Never emits a JSON `null`: fields left unset are omitted from the
`jsonld` object entirely rather than serialised as `null`.
• A REQUIRED flag omitted → `{"status":"error","reason":"bad_usage"}`,
exit 2 (bad usage, like every other verb). A required flag GIVEN but
empty (argparse cannot catch that) → `{"status":"degraded",...}`,
exit 0 — fail-open, never a traceback.
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
+3 -1
View File
@@ -34,6 +34,8 @@ case "$cmd" in
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
score)
exec "$PY" "$HERE/score.py" --store "$STORE" "$@" ;;
schema_gen)
exec "$PY" "$HERE/schema_gen.py" --store "$STORE" "$@" ;;
drift)
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
rendercheck)
@@ -52,6 +54,6 @@ case "$cmd" in
fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|forget} [flags]"}'
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|schema_gen|forget} [flags]"}'
exit 2 ;;
esac
+301
View File
@@ -0,0 +1,301 @@
#!/usr/bin/env python3
"""Deterministic JSON-LD generators for four Schema.org types. Stdlib only.
Adapted from claude-seo (github.com/AgriciDaniel/claude-seo, MIT),
schema_generate.py — rewritten to the lib/seo-data fail-open contract.
Everywhere else in this repo we AUDIT existing markup (google_seo.py
`inspect`, geo-analyzer's JSON-LD rules); this is the one verb that
GENERATES it. Reservation + potentialAction matter now that AI Mode
executes restaurant reservations; DiscussionForumPosting is a live SERP
feature; ProfilePage with sameAs/knowsAbout is the cheapest entity-graph
builder for AI citation correlation. geo-analyzer's G2 batch calls this
instead of hand-writing the markup — it only generates STRUCTURE, unknown
field VALUES stay the caller's `[À COMPLÉTER]` placeholder, never invented
here.
"""
import argparse, json
def reservation(provider, start, *, end=None, party_size=None,
reservation_id=None, reservation_for_name=None,
customer_name=None, customer_email=None,
kind="FoodEstablishmentReservation"):
"""Reservation JSON-LD block. Defaults to FoodEstablishment."""
payload = {
"@context": "https://schema.org",
"@type": kind,
"reservationStatus": "https://schema.org/ReservationConfirmed",
"provider": {"@type": "Organization", "name": provider},
"reservationFor": {
"@type": "FoodEstablishment"
if kind == "FoodEstablishmentReservation" else "Place",
"name": reservation_for_name or provider,
},
"startTime": start,
"endTime": end,
"partySize": party_size,
"reservationId": reservation_id,
}
if customer_name or customer_email:
payload["underName"] = {"@type": "Person", "name": customer_name,
"email": customer_email}
return payload
def order_action(merchant, *, order_url, name="Order online",
accepted_payment_method=None, delivery_method=None):
"""OrderAction potentialAction block. Attach to a Product/Service via
{"@type": "Product", "potentialAction": <this dict>}."""
payload = {
"@context": "https://schema.org",
"@type": "OrderAction",
"name": name,
"target": {
"@type": "EntryPoint",
"urlTemplate": order_url,
"inLanguage": "en-US",
"actionPlatform": [
"https://schema.org/DesktopWebPlatform",
"https://schema.org/MobileWebPlatform",
],
},
"deliveryMethod": delivery_method or [
"https://schema.org/OnSitePickup",
"https://schema.org/ParcelService",
],
"priceSpecification": {
"@type": "PriceSpecification",
"eligibleTransactionVolume": {
"@type": "PriceSpecification",
"minPrice": 0,
"priceCurrency": "USD",
},
},
"merchant": {"@type": "Organization", "name": merchant},
}
if accepted_payment_method:
payload["acceptedPaymentMethod"] = [
{"@type": "PaymentMethod", "name": m}
for m in accepted_payment_method
]
return payload
def discussion(headline, author, *, url, date_published, text=None,
date_modified=None, interaction_count=None,
comment_count=None):
"""DiscussionForumPosting JSON-LD block."""
payload = {
"@context": "https://schema.org",
"@type": "DiscussionForumPosting",
"headline": headline,
"author": {"@type": "Person", "name": author},
"datePublished": date_published,
"dateModified": date_modified,
"url": url,
"mainEntityOfPage": {"@type": "WebPage", "@id": url},
"text": text,
"commentCount": comment_count,
}
if interaction_count:
payload["interactionStatistic"] = [
{"@type": "InteractionCounter",
"interactionType": "https://schema.org/%s" % k,
"userInteractionCount": v}
for k, v in interaction_count.items()
]
return payload
def profile(name, *, url, description=None, same_as=None, knows_about=None,
works_for=None, image=None, job_title=None):
"""ProfilePage JSON-LD block. sameAs + knowsAbout is the entity-graph
helper for AI citation correlation — Wikipedia/GitHub/LinkedIn/ORCID
URLs in sameAs disambiguate the person across knowledge graphs."""
person = {
"@type": "Person",
"name": name,
"url": url,
"description": description,
"sameAs": list(same_as) if same_as else None,
"knowsAbout": list(knows_about) if knows_about else None,
"worksFor": {"@type": "Organization", "name": works_for}
if works_for else None,
"image": image,
"jobTitle": job_title,
}
return {"@context": "https://schema.org", "@type": "ProfilePage",
"mainEntity": person, "url": url}
def _strip_nones(value):
"""Recursively drop dict keys AND list elements whose value is None —
the emitted JSON-LD must never contain a null."""
if isinstance(value, dict):
return {k: _strip_nones(v) for k, v in value.items() if v is not None}
if isinstance(value, list):
return [_strip_nones(v) for v in value if v is not None]
return value
def _need(value, field):
"""Raise on a schema-required field that is present but empty — the
case argparse's `required=True` cannot catch (an empty string is a
given flag, not a missing one)."""
if value is None or not str(value).strip():
raise ValueError("missing required field: %s" % field)
return value
def _generate(kind, args):
"""Route to the matching generator, enforcing schema-required fields."""
if kind == "reservation":
return reservation(
_need(args.provider, "provider"), _need(args.start, "start"),
end=args.end, party_size=args.party_size,
reservation_id=args.reservation_id,
reservation_for_name=args.reservation_for_name,
customer_name=args.customer_name,
customer_email=args.customer_email, kind=args.reservation_kind,
)
if kind == "order":
return order_action(
_need(args.merchant, "merchant"),
order_url=_need(args.order_url, "order_url"), name=args.name,
accepted_payment_method=args.accepted_payment_method,
delivery_method=args.delivery_method,
)
if kind == "discussion":
interaction = {"LikeAction": args.likes} if args.likes else None
return discussion(
_need(args.headline, "headline"), _need(args.author, "author"),
url=_need(args.url, "url"),
date_published=_need(args.date_published, "date_published"),
text=args.text, date_modified=args.date_modified,
interaction_count=interaction, comment_count=args.comment_count,
)
if kind == "profile":
return profile(
_need(args.name, "name"), url=_need(args.url, "url"),
description=args.description, same_as=args.same_as,
knows_about=args.knows_about, works_for=args.works_for,
image=args.image, job_title=args.job_title,
)
raise ValueError("unknown kind: %r" % kind) # pragma: no cover — argparse
def _envelope(payload, script_tag):
cleaned = _strip_nones(payload)
out = {"status": "ok", "source": "schema_gen",
"type": cleaned.get("@type"), "jsonld": cleaned}
if script_tag:
pretty = json.dumps(cleaned, indent=2, ensure_ascii=False)
out["script"] = ('<script type="application/ld+json">\n%s\n</script>'
% pretty)
return out
def _script_tag_parent():
"""`--script-tag` as a shared parent parser, so it is valid on every
subcommand — `fetch.sh schema_gen <type> [flags]` puts the type FIRST,
and argparse only accepts a flag after a subcommand token if that flag
was declared on the subparser, not the top-level one."""
parent = argparse.ArgumentParser(add_help=False)
parent.add_argument(
"--script-tag", action="store_true",
help="Wrap jsonld in <script type=application/ld+json>.",
)
return parent
def _add_reservation_args(sub, parents):
p = sub.add_parser("reservation", parents=parents,
help="FoodEstablishmentReservation et al.")
p.add_argument("--provider", required=True)
p.add_argument("--start", required=True, help="ISO 8601 startTime.")
p.add_argument("--end")
p.add_argument("--party-size", type=int)
p.add_argument("--reservation-id")
p.add_argument("--reservation-for-name")
p.add_argument("--customer-name")
p.add_argument("--customer-email")
p.add_argument(
"--reservation-kind", dest="reservation_kind",
default="FoodEstablishmentReservation",
choices=(
"FoodEstablishmentReservation", "LodgingReservation",
"RentalCarReservation", "TaxiReservation", "EventReservation",
"TrainReservation", "FlightReservation",
),
)
def _add_order_args(sub, parents):
p = sub.add_parser("order", parents=parents,
help="OrderAction (potentialAction).")
p.add_argument("--merchant", required=True)
p.add_argument("--order-url", required=True)
p.add_argument("--name", default="Order online")
p.add_argument("--accepted-payment-method", nargs="*", default=None)
p.add_argument("--delivery-method", nargs="*", default=None)
def _add_discussion_args(sub, parents):
p = sub.add_parser("discussion", parents=parents,
help="DiscussionForumPosting.")
p.add_argument("--headline", required=True)
p.add_argument("--author", required=True)
p.add_argument("--url", required=True)
p.add_argument("--date", dest="date_published", required=True)
p.add_argument("--text")
p.add_argument("--date-modified")
p.add_argument("--comment-count", type=int)
p.add_argument("--likes", type=int, default=None,
help="LikeAction count (interactionStatistic).")
def _add_profile_args(sub, parents):
p = sub.add_parser("profile", parents=parents,
help="ProfilePage with sameAs / knowsAbout.")
p.add_argument("--name", required=True)
p.add_argument("--url", required=True)
p.add_argument("--description")
p.add_argument("--same-as", nargs="*", default=None)
p.add_argument("--knows-about", nargs="*", default=None)
p.add_argument("--works-for")
p.add_argument("--image")
p.add_argument("--job-title")
def _build_parser():
p = argparse.ArgumentParser(
description="Schema.org JSON-LD generators (stdlib, deterministic)."
)
p.add_argument("--store", default=None) # accepted+ignored (dispatch)
sub = p.add_subparsers(dest="kind", required=True)
parents = [_script_tag_parent()]
_add_reservation_args(sub, parents)
_add_order_args(sub, parents)
_add_discussion_args(sub, parents)
_add_profile_args(sub, parents)
return p
def _cli():
try:
args = _build_parser().parse_args()
payload = _generate(args.kind, args)
print(json.dumps(_envelope(payload, args.script_tag), indent=2))
except SystemExit as e:
if e.code not in (0, None):
print(json.dumps({"status": "error", "reason": "bad_usage"}))
raise
except Exception as e:
# Fail-open: a missing required field or any other unexpected error
# is a normal outcome here, never a traceback or empty stdout.
print(json.dumps({"status": "degraded", "reason": str(e)}))
if __name__ == "__main__":
_cli()
+59
View File
@@ -233,6 +233,63 @@ NCHG="$(printf '%s' "$D2" | python3 -c 'import sys,json; print(len(json.load(sys
|| no "reworded title is a change, not a regression" "got $NCHG"
rm -rf "$DH"
echo "── schema_gen ──"
SG() { python3 "$SD/schema_gen.py" "$@"; }
RES="$(SG reservation --provider "Chez X" --start "2026-08-01T19:00")"
has "reservation ok" "$RES" '"status": "ok"'
has "reservation type surfaced" "$RES" '"type": "FoodEstablishmentReservation"'
has "jsonld has @context" "$RES" '"@context": "https://schema.org"'
has "reservation keeps provider" "$RES" 'Chez X'
has "reservation keeps start" "$RES" '2026-08-01T19:00'
PROF="$(SG profile --name "Jane Doe" --url https://ex.com/about)"
has "profile ok" "$PROF" '"status": "ok"'
has "profile type surfaced" "$PROF" '"type": "ProfilePage"'
ORD="$(SG order --merchant "Acme" --order-url https://ex.com/order)"
has "order ok" "$ORD" '"status": "ok"'
has "order type surfaced" "$ORD" '"type": "OrderAction"'
DISC="$(SG discussion --headline "Q" --author "Jo" --url https://ex.com/t/1 \
--date 2026-05-01T00:00:00Z)"
has "discussion ok" "$DISC" '"status": "ok"'
has "discussion type surfaced" "$DISC" '"type": "DiscussionForumPosting"'
# argparse required=True catches an OMITTED flag → bad usage, exit 2
BADRES="$(SG reservation --start 2026-08-01T19:00 2>/dev/null)"; BADRC=$?
hasnt "missing --provider is not ok" "$BADRES" '"status": "ok"'
[ "$BADRC" = "2" ] && ok "missing --provider exit 2" \
|| no "missing --provider exit 2" "got $BADRC"
# a required field argparse ALLOWS through (flag given, value empty) must
# still fail open — degraded, not a crash, exit 0
EMPTYRES="$(SG reservation --provider "" --start 2026-08-01T19:00)"; EMPTYRC=$?
hasnt "empty --provider is not ok" "$EMPTYRES" '"status": "ok"'
has "empty --provider degrades" "$EMPTYRES" '"status": "degraded"'
[ "$EMPTYRC" = "0" ] && ok "empty --provider exit 0" \
|| no "empty --provider exit 0" "got $EMPTYRC"
# --script-tag must work AFTER the type, matching `fetch.sh schema_gen
# <type> [flags]` — the shape the dispatcher actually calls it with. The
# envelope is JSON, so the `script` field's own quotes are backslash-escaped
# in the raw stdout — decode it to check the LITERAL wrapper string.
SCRIPT="$(SG profile --name "Jane Doe" --url https://ex.com/about --script-tag)"
SCRIPT_TAG="$(printf '%s' "$SCRIPT" | \
python3 -c 'import sys,json; print(json.load(sys.stdin)["script"])')"
has "script-tag wraps output" "$SCRIPT_TAG" '<script type="application/ld+json">'
# an omitted optional field must never surface as a JSON null
hasnt "no null ever emitted" "$RES" 'null'
# stdlib ONLY — no requests/httpx/bs4/any third-party import
IMPORTS="$(grep -E '^(import|from) ' "$SD/schema_gen.py")"
if printf '%s' "$IMPORTS" | grep -qiE 'requests|httpx|bs4'; then
no "schema_gen stdlib only" "third-party import found: $IMPORTS"
else
ok "schema_gen stdlib only"
fi
# dispatch wiring: --store precedes the type (fetch.sh's own convention),
# --script-tag comes after it (the caller's convention) — both must work
# through the real fetch.sh entrypoint, not just the bare script
FSG="$(SEO_DATA_ENV_FILE=/dev/null SEO_DATA_STORE=/nonexistent bash "$SD/fetch.sh" \
schema_gen reservation --provider "Chez X" --start 2026-08-01T19:00 --script-tag)"
has "fetch dispatches schema_gen" "$FSG" '"status": "ok"'
FSG_TAG="$(printf '%s' "$FSG" | \
python3 -c 'import sys,json; print(json.load(sys.stdin)["script"])')"
has "fetch schema_gen script-tag" "$FSG_TAG" '<script type="application/ld+json">'
echo "── fetch.sh ──"
FETCH="$SD/fetch.sh"
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
@@ -358,6 +415,7 @@ tf "analyzer calls fetch crux" "$REPO/agents/seo-analyzer.md" "fetch.sh crux"
tf "analyzer calls fetch queries" "$REPO/agents/seo-analyzer.md" "fetch.sh queries"
tf "analyzer gsc subsection" "$REPO/agents/seo-analyzer.md" "Performance GSC"
tf "catalog gsc oauth entry" "$REPO/agents/resources/automation-catalog.md" "make seo-connect"
tf "geo-analyzer wires schema_gen" "$REPO/agents/geo-analyzer.md" "fetch.sh schema_gen"
echo "── account-mgmt locks ──"
tf "skill routes account verbs" "$REPO/skills/seo/SKILL.md" "forget --all"
@@ -370,6 +428,7 @@ tf "readme documents fetch.sh" "$REPO/lib/seo-data/README.md" "fetch.sh"
tf "readme documents seo-connect" "$REPO/lib/seo-data/README.md" "make seo-connect"
tf "readme documents forget" "$REPO/lib/seo-data/README.md" "forget --all"
tf "readme revocation note" "$REPO/lib/seo-data/README.md" "myaccount.google.com/permissions"
tf "readme documents schema_gen" "$REPO/lib/seo-data/README.md" "schema_gen"
echo ""
echo "seo-data engine: $PASS pass, $FAIL fail"