Commit Graph
9 Commits
Author SHA1 Message Date
Bastien Chanot cfdd89e73b feat(seo-data): schema_gen verb — generate JSON-LD, not just audit it
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.

fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.

Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).

geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.

Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 14:30:31 +02:00
Bastien Chanot 4818c6116f feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.

The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.

Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.

Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
  off-page) both mandate excluding an axis and renormalising the rest. Both
  left that arithmetic to the model. Now the engine does it and refuses to let
  N/A behave like a zero — verified: all-20 axes with two N/A still yields
  global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
  single page de-escalates). A defect on 1 of 12 pages is not the defect on
  12 of 12, and flattening the two is part of what made the old numbers move.

Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.

Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 13:29:14 +02:00
Bastien Chanot f69cfc5cb4 feat(seo-data): H2 — drift baseline; regressions vs changes, not prose
seo-analyzer.md:1365 keeps history as "date + score + key changes" — prose the
LLM writes about its own previous prose. Lossy, unreproducible, and
machine-uncomparable, so "the redesign silently dropped 40 canonicals" is
invisible unless someone happens to notice.

drift snapshots title/description/canonical/robots/h1_count/jsonld_types per
URL and diffs them. Stdlib only, no auth.

The classification IS the feature: LOSING a signal is a regression, CHANGING
one is a change that may well be intended. The engine says which kind; the
agent judges. A reworded title is not an alert; an evaporated canonical is.

Runs over the WHOLE sitemap, never a sample — caught while designing: a drift
computed over a sample that changes between runs compares nothing.

NOT rank tracking. That is the common misread of this same feature elsewhere;
positions come from GSC `queries`. This is on-page regression detection.

Also caught in my own draft before testing: _capture reused
sm._mock("page.html"), the exact single-fixture flaw I had already fixed in
linkgraph — one fixture cannot express a multi-page snapshot, every URL would
read identical. Now pages.json, same convention.

Proved on a planted failure rather than a happy path — two clean sites would
look identical to a detector that always returns []:
  v1 -> v2: canonical lost on /a, h1 + jsonld lost on /, title reworded,
  /gone removed, /neuve added
  → 3 regressions, 1 change, gone/new both detected, title correctly NOT a
    regression.

Store is ~/.claude/seo-data/drift/<host>.json, 0700, written via os.replace so
a crash never leaves a half-written baseline; a corrupt store degrades to
"first run" instead of killing the audit.

Verified: seo-data 144 -> 155 pass, 0 fail; full suite green.
2026-07-17 13:25:45 +02:00
Bastien Chanot 20d3082542 feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright
Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
2026-07-17 12:44:43 +02:00
Bastien Chanot fe41986be9 feat(seo-data): C3 — internal link graph; orphans + click depth, measured
seo-analyzer.md:613 asks "Every important page reachable within 3 clicks?"
and :616 asks "Orphan pages (no inbound internal links)?". Neither ever had a
command — same shape as the sameAs check before W3. This is that command.

My earlier reservation ("costs a lot of network") was wrong and the
measurement killed it: 24 pages in 2.7s, 86 in 3.8s. Cheap enough to always
run on FULL.

EXHAUSTIVE OR NOTHING is the design constraint, not a nicety. Orphans cannot
be sampled: proving a page has no inbound link means having read every other
page. So when the crawl is capped or any page fails, orphans are WITHHELD —
`orphans_withheld: true` and no list. A false orphan ("page X has no inbound
links" when it does) sends a client fixing what is not broken; that is the
worst finding this tool could emit. The cap does not degrade the result, it
invalidates it.

SPA refusal: on a client-rendered site the links are not in the HTML and
every page reads as orphaned. That is catastrophic, so an empty graph returns
degraded/no_links_in_html instead of a full false-positive list. No JS
rendering by design — that is the R1/R2 arbitration, not something to smuggle
in here.

Verified against BOTH live sites and against a planted failure, because two
clean results are not evidence a detector detects:
- native PHP: 24 pages, 335 links, depth 2, 0 orphans
- Astro: 86 pages, 2015 links, depth 2, 0 orphans
- fixture with a planted orphan + a 4-click chain: both found. Filters proven
  on real shapes seen live — /css/main.css?v=1778157313, #anchors, mailto:,
  tel:, external hosts, .png. /b/ in markup vs /b in sitemap unify to one node
  rather than a phantom orphan pair.

Fixed a flaw in my own mock while writing that test: a single page.html
fixture cannot express a GRAPH (every node gets identical links), so the mock
is now pages.json = {url: html}.

Verified: seo-data 122 -> 136 pass, 0 fail; full suite green; shellcheck +
py_compile clean.
2026-07-17 12:32:42 +02:00
Bastien Chanot 3a15643c2c feat(seo-data): C2 — cannibalisation from Google's own data, one param away
The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.

CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.

  fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
  impressions, strongest page first inside each. Same auth, same quota family,
  no new scope. `capped` reports a full row window rather than presenting a
  truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.

Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.

Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.

30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.

The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.

Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 11:55:00 +02:00
Bastien Chanot 2de58faa38 feat(seo-data): C1b — sitemap verb, the denominator COVERAGE never had
I5 made a COVERAGE line mandatory in STEP 9 and told the agent to "count the
URLs in sitemap.xml" without giving it a command. STEP 4 only ever did
`curl … | head -50` — a preview, not a count. This closes that.

fetch.sh sitemap --url … → {count, urls[], index, dropped}. Stdlib only
(urllib + xml.etree + gzip): no auth, no Google, no venv, so it runs wherever
the mock/degrade paths run. Follows <sitemapindex> one level, dedupes, strips
whitespace, handles .xml.gz. Every cap REPORTS what it cut (children_skipped,
truncated) rather than truncating silently — same rule as COVERAGE itself.

PLAN CORRECTION: the proposal said the verb would "validate each URL via the
H1 guard". Wrong. urllib fetches these, so nothing here reaches a shell and
there is no injection surface to guard. The guard belongs at the point of
use, where seo-analyzer interpolates a URL into curl — which is the contract
the sameAs check already established. A second copy of url-guard here would
only drift from the first. The module carries a garbage filter, named as such.

SECURITY: the security-guidance hook asked for defusedxml. Taken seriously,
not obeyed — it would drag a venv into a module whose whole point is being
stdlib-only. Split the threat instead: xml.etree does NOT expand external
entities (XXE is not the vector), but it IS billion-laughs-vulnerable, and
the 20 MB read ceiling bounds the input, not the expansion. A sitemap NEVER
has a DTD — sitemaps.org is <?xml?> then <urlset xmlns=> — so any
doctype/entity is refused BEFORE parsing, with its own reason
(unsafe_xml_dtd, distinct from parse_failed: it is a finding, not a glitch).
Refusing the construct beats depending on parser internals. Fixture is a real
billion-laughs payload.

Verified against the live target, not just fixtures: zenquality's sitemap
returns count=86, dropped=0, matching `grep -c '<loc>'` on the raw XML
exactly. Dead URL → {"status":"degraded","reason":"fetch_failed"}, exit 0.
seo-data 95 -> 110 pass, 0 fail; full suite green; shellcheck + py_compile
clean.

Note: no config-edit sentinel was needed after all — config-protection guards
lib/tests, not lib/seo-data. I posted one, found it uncommitted-and-unconsumed
afterwards, and removed it rather than leave an open one-shot gate lying
around. Worth knowing: seo-data.test.sh is 110 assertions and is NOT covered
by that hook, while lib/tests/*.test.sh is.
2026-07-17 11:28:00 +02:00
Bastien Chanot 8bf7459566 feat(seo): account-management verbs (connect/accounts/forget) + connect.sh wrapper
tokenstore remove/clear, fetch.sh forget dispatch, and a connect.sh wrapper
that sources ~/.claude/.env internally and runs from any project. /seo now
routes connect|accounts|forget before the audit flow; Makefile seo-connect
delegates to the wrapper. Labels are guarded to shell-safe ASCII (POSIX case,
whole-string, C-locale) as defense-in-depth; forget output states local
removal is not a Google-side revocation.
2026-07-10 12:38:32 +02:00
Bastien Chanot c3a504fbbf feat(seo-data): fetch.sh entrypoint with venv/system fallback and redaction 2026-07-10 01:32:48 +02:00