feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright

Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
This commit is contained in:
Bastien Chanot
2026-07-17 12:44:43 +02:00
parent fe41986be9
commit 20d3082542
8 changed files with 225 additions and 1 deletions
+11
View File
@@ -515,6 +515,17 @@ PRIORITY ACTIONS : <top 3-5>
## STEP 8 — CONTENT SHAPE FOR AI `[both]`
**Rendering gate first (R2).** `bash ~/.claude/lib/seo-data/fetch.sh
rendercheck --url "https://$DOMAIN/"`. Verdict `client-rendered` → Content
Shape is `N/A — content not in served HTML`, excluded from the weighted
global, never scored zero. And say the thing that actually matters here: AI
crawlers are **worse** at JS than Googlebot is. GPTBot, PerplexityBot and
ClaudeBot fetch HTML and largely do not execute it, so a client-rendered site
is not just unauditable by us — it is close to invisible to the engines this
whole audit targets. That is a §0 alert and the top user action (SSR/SSG),
not a schema tweak.
Site-wide axes (crawler policy, llms.txt) are unaffected: those are files.
Load: `~/.claude/agents/resources/content-shape-for-ai.md`
Sample 5-10 key pages (homepage + top service/blog pages). For each: