From 20d30825422b09ddcbdf96d220e28eff1e91ae58 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 12:44:43 +0200 Subject: [PATCH] =?UTF-8?q?feat(seo-data,seo,geo):=20R2=20=E2=80=94=20refu?= =?UTF-8?q?se=20to=20score=20what=20JS=20paints;=20no=20Playwright?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Arbitrated (user): honest refusal on SPA, no headless browser. STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag compensated — it does not exist. Seventh subagent claim this branch has had to disprove.) So on a client-rendered site the FULL audit curls an empty shell, every meta/H1/JSON-LD check reports "missing", and the agent emits a page of false findings — plus a bundle that would "fix" tags which already exist. rendercheck reads the verdict from what the server SENT. package.json cannot tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only. The refusal is the point: - client-rendered → On-page is N/A, excluded from the weighted global, NOT scored zero. A zero says "your on-page is bad"; N/A says "we could not see it". Only one is true, and /client-handover gates on this number. - No bundle item may come from a live on-page check on such a site. - The report still says what IS auditable (robots, sitemap, headers, config, CrUX field data — real users, hydration included — GSC, legal, images) rather than returning an empty verdict. - geo refuses Content Shape the same way, and states the sharper fact: AI crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot fetch HTML and largely do not execute it, so a client-rendered site is not merely unauditable by us — it is near-invisible to the engines this audit exists to serve. §0 alert + SSR/SSG as the top user action. Script/style text is not page text: a React shell with a fat inline window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle reads as a rich page — the detector would fail exactly where it matters. Verified on both extremes, not just the happy path: zenquality 7650 chars/1 h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a Vite/React shell fixture → client-rendered, 7/0/0, warned. seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile clean. --- agents/geo-analyzer.md | 11 +++ agents/seo-analyzer.md | 50 +++++++++++++ lib/seo-data/README.md | 23 ++++++ lib/seo-data/fetch.sh | 4 +- lib/seo-data/fixtures-spa/page.html | 8 +++ lib/seo-data/fixtures-ssr/page.html | 8 +++ lib/seo-data/render_check.py | 106 ++++++++++++++++++++++++++++ lib/seo-data/seo-data.test.sh | 16 +++++ 8 files changed, 225 insertions(+), 1 deletion(-) create mode 100644 lib/seo-data/fixtures-spa/page.html create mode 100644 lib/seo-data/fixtures-ssr/page.html create mode 100644 lib/seo-data/render_check.py diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index b9d870f..66f845c 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -515,6 +515,17 @@ PRIORITY ACTIONS : ## STEP 8 — CONTENT SHAPE FOR AI `[both]` +**Rendering gate first (R2).** `bash ~/.claude/lib/seo-data/fetch.sh +rendercheck --url "https://$DOMAIN/"`. Verdict `client-rendered` → Content +Shape is `N/A — content not in served HTML`, excluded from the weighted +global, never scored zero. And say the thing that actually matters here: AI +crawlers are **worse** at JS than Googlebot is. GPTBot, PerplexityBot and +ClaudeBot fetch HTML and largely do not execute it, so a client-rendered site +is not just unauditable by us — it is close to invisible to the engines this +whole audit targets. That is a §0 alert and the top user action (SSR/SSG), +not a schema tweak. +Site-wide axes (crawler policy, llms.txt) are unaffected: those are files. + Load: `~/.claude/agents/resources/content-shape-for-ai.md` Sample 5-10 key pages (homepage + top service/blog pages). For each: diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 7477f96..6f315ff 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -472,6 +472,48 @@ Fetch rendered HTML. Extract and analyze: ## STEP 5 — ON-PAGE AUDIT `[both]` +### Rendering gate — run this BEFORE anything else in STEP 5 (R2) + +```bash +bash ~/.claude/lib/seo-data/fetch.sh rendercheck --url "https://$DOMAIN/" +``` + +STEP 2 has always recorded `RENDERING: SSR/SSG/SPA/hybrid` and nothing ever +acted on it. This is the rule that does. The verdict comes from what the +server actually sent, not from reading package.json — a React SPA and a +Next.js SSR app are indistinguishable there. + +**`verdict: client-rendered` → REFUSE to score the On-page axis.** Do not +score it low. Do not score it at all: +- On-page → `N/A — content not in served HTML (client-rendered)`. Redistribute + nothing; a missing axis is not a zero. +- Every curl-based meta/H1/JSON-LD check would report "missing" against a site + that may be perfectly correct once hydrated. Those are FALSE findings, and + a bundle built on them would "fix" meta tags that already exist. +- **No bundle item may come from a live on-page check on this site.** Source + greps still apply — the JSX carries the tags — but you cannot tell which + route renders what, so treat them as inventory, not as per-page findings. +- `linkgraph` will refuse too (`no_links_in_html`) — the same blindness. Do + not work around either refusal. + +Still fully auditable, and worth saying so rather than returning an empty +report: robots.txt, sitemap.xml, HTTP headers, redirects, `.htaccess` / +framework config, CWV via CrUX (field data is real-user, hydration included), +GSC queries + index coverage, legal pages, image weights. + +**`verdict: partial`** → shell plus an SSR'd head, or a genuinely thin page. +Score what is present, name what is not, and say which of the two you think +it is. + +**§0 line, mandatory when not server-rendered:** +`Rendering: client-rendered — On-page NOT scored (content absent from served +HTML). Global score excludes it. Fix: SSR/SSG (CLAUDE.md: public sites are +never SPAs).` + +This is the honest half of the R1/R2 call: we do not render JS (no Playwright, +no Chromium), so we do not pretend to see what JS paints. Refusing is the +finding. + **Record the denominator BEFORE sampling.** This step samples; the report says "audit". On a 500-page site a 12-page sample is 2.4% — the On-page score is an extrapolation from it, and the reader cannot know unless you print it. @@ -885,6 +927,14 @@ owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden Name what you saw. An omission has to stay legible — the same reason COVERAGE is mandatory in STEP 9. +**On-page axis note (R2).** `rendercheck` verdict `client-rendered` → this +axis is `N/A — content not in served HTML`, excluded from the weighted global, +NOT scored zero. A zero says "your on-page is bad"; N/A says "we could not +see it", and only one of those is true. Renormalise the remaining weights over +the axes actually scored and say so on the SEO GLOBAL line. The code ceiling +must state that no code fix raises an axis we did not measure — the unlock is +SSR/SSG, and that is a user action, not a bundle item. + **Off-page axis note (I1).** Score ONLY the unlinked brand mentions gathered in STEP 6 (`web_search "" -site:`). Backlink profile and domain authority have NO data source here — no index, diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index 276a154..dfbbbfa 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -146,6 +146,29 @@ fetch.sh sitemap --url https://ex.com/sitemap.xml internals AND keeps this stdlib-only; defusedxml would drag in a venv for a document type that has no legitimate DTD. +fetch.sh rendercheck --url https://ex.com/ + → {"status":"ok","verdict":"server-rendered"|"client-rendered"|"partial", + "body_text_chars":7650,"h1_in_html":1,"jsonld_in_html":9, + "meta_description_in_html":true,"html_bytes":132447, + "warning":"…"} # warning only when not server-rendered + + R2, the honest half of the SPA call. seo-analyzer has always recorded + `RENDERING: SSR/SSG/SPA` and never acted on it; this is the signal it acts + on. Verdict comes from what the server SENT — package.json cannot tell a + React SPA from a Next.js SSR app. + • client-rendered → the agent REFUSES to score On-page (N/A, not zero: a + zero says "your on-page is bad", N/A says "we could not see it"). Every + curl-based meta/H1/JSON-LD check would report "missing" against a site + that is fine once hydrated — false findings, and a bundle that "fixes" + tags which already exist. + • Does NOT render JS. No Playwright, no Chromium, no venv. Refusing IS the + finding. + • Script/style text is not page text: measured 7 chars on a React shell + whose inline window.__INITIAL_STATE__ is large. Without that, a 200 KB + bundle reads as a rich page. + • Measured 2026-07-17: zenquality 7650 chars/1 h1/9 jsonld and + lavageangels356 13973/1/1 → server-rendered; a Vite shell → 7/0/0. + fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500] → {"status":"ok","source":"linkgraph","pages_crawled":86,"pages_failed":0, "total_internal_links":2015,"capped":false,"max_depth":2, diff --git a/lib/seo-data/fetch.sh b/lib/seo-data/fetch.sh index 3f29dfa..7ba61fc 100644 --- a/lib/seo-data/fetch.sh +++ b/lib/seo-data/fetch.sh @@ -32,6 +32,8 @@ case "$cmd" in # No auth, no Google: stdlib-only, runs even without the venv. sitemap) exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;; + rendercheck) + exec "$PY" "$HERE/render_check.py" --store "$STORE" "$@" ;; linkgraph) exec "$PY" "$HERE/linkgraph.py" --store "$STORE" "$@" ;; forget) @@ -46,6 +48,6 @@ case "$cmd" in fi echo '{"status":"error","reason":"usage: fetch.sh forget {--label