feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright

Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
This commit is contained in:
Bastien Chanot
2026-07-17 12:44:43 +02:00
parent fe41986be9
commit 20d3082542
8 changed files with 225 additions and 1 deletions
+11
View File
@@ -515,6 +515,17 @@ PRIORITY ACTIONS : <top 3-5>
## STEP 8 — CONTENT SHAPE FOR AI `[both]` ## STEP 8 — CONTENT SHAPE FOR AI `[both]`
**Rendering gate first (R2).** `bash ~/.claude/lib/seo-data/fetch.sh
rendercheck --url "https://$DOMAIN/"`. Verdict `client-rendered` → Content
Shape is `N/A — content not in served HTML`, excluded from the weighted
global, never scored zero. And say the thing that actually matters here: AI
crawlers are **worse** at JS than Googlebot is. GPTBot, PerplexityBot and
ClaudeBot fetch HTML and largely do not execute it, so a client-rendered site
is not just unauditable by us — it is close to invisible to the engines this
whole audit targets. That is a §0 alert and the top user action (SSR/SSG),
not a schema tweak.
Site-wide axes (crawler policy, llms.txt) are unaffected: those are files.
Load: `~/.claude/agents/resources/content-shape-for-ai.md` Load: `~/.claude/agents/resources/content-shape-for-ai.md`
Sample 5-10 key pages (homepage + top service/blog pages). For each: Sample 5-10 key pages (homepage + top service/blog pages). For each:
+50
View File
@@ -472,6 +472,48 @@ Fetch rendered HTML. Extract and analyze:
## STEP 5 — ON-PAGE AUDIT `[both]` ## STEP 5 — ON-PAGE AUDIT `[both]`
### Rendering gate — run this BEFORE anything else in STEP 5 (R2)
```bash
bash ~/.claude/lib/seo-data/fetch.sh rendercheck --url "https://$DOMAIN/"
```
STEP 2 has always recorded `RENDERING: SSR/SSG/SPA/hybrid` and nothing ever
acted on it. This is the rule that does. The verdict comes from what the
server actually sent, not from reading package.json — a React SPA and a
Next.js SSR app are indistinguishable there.
**`verdict: client-rendered` → REFUSE to score the On-page axis.** Do not
score it low. Do not score it at all:
- On-page → `N/A — content not in served HTML (client-rendered)`. Redistribute
nothing; a missing axis is not a zero.
- Every curl-based meta/H1/JSON-LD check would report "missing" against a site
that may be perfectly correct once hydrated. Those are FALSE findings, and
a bundle built on them would "fix" meta tags that already exist.
- **No bundle item may come from a live on-page check on this site.** Source
greps still apply — the JSX carries the tags — but you cannot tell which
route renders what, so treat them as inventory, not as per-page findings.
- `linkgraph` will refuse too (`no_links_in_html`) — the same blindness. Do
not work around either refusal.
Still fully auditable, and worth saying so rather than returning an empty
report: robots.txt, sitemap.xml, HTTP headers, redirects, `.htaccess` /
framework config, CWV via CrUX (field data is real-user, hydration included),
GSC queries + index coverage, legal pages, image weights.
**`verdict: partial`** → shell plus an SSR'd head, or a genuinely thin page.
Score what is present, name what is not, and say which of the two you think
it is.
**§0 line, mandatory when not server-rendered:**
`Rendering: client-rendered — On-page NOT scored (content absent from served
HTML). Global score excludes it. Fix: SSR/SSG (CLAUDE.md: public sites are
never SPAs).`
This is the honest half of the R1/R2 call: we do not render JS (no Playwright,
no Chromium), so we do not pretend to see what JS paints. Refusing is the
finding.
**Record the denominator BEFORE sampling.** This step samples; the report **Record the denominator BEFORE sampling.** This step samples; the report
says "audit". On a 500-page site a 12-page sample is 2.4% — the On-page score says "audit". On a 500-page site a 12-page sample is 2.4% — the On-page score
is an extrapolation from it, and the reader cannot know unless you print it. is an extrapolation from it, and the reader cannot know unless you print it.
@@ -885,6 +927,14 @@ owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden
Name what you saw. An omission has to stay legible — the same reason Name what you saw. An omission has to stay legible — the same reason
COVERAGE is mandatory in STEP 9. COVERAGE is mandatory in STEP 9.
**On-page axis note (R2).** `rendercheck` verdict `client-rendered` → this
axis is `N/A — content not in served HTML`, excluded from the weighted global,
NOT scored zero. A zero says "your on-page is bad"; N/A says "we could not
see it", and only one of those is true. Renormalise the remaining weights over
the axes actually scored and say so on the SEO GLOBAL line. The code ceiling
must state that no code fix raises an axis we did not measure — the unlock is
SSR/SSG, and that is a user action, not a bundle item.
**Off-page axis note (I1).** Score ONLY the unlinked brand mentions **Off-page axis note (I1).** Score ONLY the unlinked brand mentions
gathered in STEP 6 (`web_search "<business-name>" -site:<domain>`). gathered in STEP 6 (`web_search "<business-name>" -site:<domain>`).
Backlink profile and domain authority have NO data source here — no index, Backlink profile and domain authority have NO data source here — no index,
+23
View File
@@ -146,6 +146,29 @@ fetch.sh sitemap --url https://ex.com/sitemap.xml
internals AND keeps this stdlib-only; defusedxml would drag in a venv for internals AND keeps this stdlib-only; defusedxml would drag in a venv for
a document type that has no legitimate DTD. a document type that has no legitimate DTD.
fetch.sh rendercheck --url https://ex.com/
→ {"status":"ok","verdict":"server-rendered"|"client-rendered"|"partial",
"body_text_chars":7650,"h1_in_html":1,"jsonld_in_html":9,
"meta_description_in_html":true,"html_bytes":132447,
"warning":"…"} # warning only when not server-rendered
R2, the honest half of the SPA call. seo-analyzer has always recorded
`RENDERING: SSR/SSG/SPA` and never acted on it; this is the signal it acts
on. Verdict comes from what the server SENT — package.json cannot tell a
React SPA from a Next.js SSR app.
• client-rendered → the agent REFUSES to score On-page (N/A, not zero: a
zero says "your on-page is bad", N/A says "we could not see it"). Every
curl-based meta/H1/JSON-LD check would report "missing" against a site
that is fine once hydrated — false findings, and a bundle that "fixes"
tags which already exist.
• Does NOT render JS. No Playwright, no Chromium, no venv. Refusing IS the
finding.
• Script/style text is not page text: measured 7 chars on a React shell
whose inline window.__INITIAL_STATE__ is large. Without that, a 200 KB
bundle reads as a rich page.
• Measured 2026-07-17: zenquality 7650 chars/1 h1/9 jsonld and
lavageangels356 13973/1/1 → server-rendered; a Vite shell → 7/0/0.
fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500] fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500]
→ {"status":"ok","source":"linkgraph","pages_crawled":86,"pages_failed":0, → {"status":"ok","source":"linkgraph","pages_crawled":86,"pages_failed":0,
"total_internal_links":2015,"capped":false,"max_depth":2, "total_internal_links":2015,"capped":false,"max_depth":2,
+3 -1
View File
@@ -32,6 +32,8 @@ case "$cmd" in
# No auth, no Google: stdlib-only, runs even without the venv. # No auth, no Google: stdlib-only, runs even without the venv.
sitemap) sitemap)
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;; exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
rendercheck)
exec "$PY" "$HERE/render_check.py" --store "$STORE" "$@" ;;
linkgraph) linkgraph)
exec "$PY" "$HERE/linkgraph.py" --store "$STORE" "$@" ;; exec "$PY" "$HERE/linkgraph.py" --store "$STORE" "$@" ;;
forget) forget)
@@ -46,6 +48,6 @@ case "$cmd" in
fi fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}' echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;; exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|linkgraph|forget} [flags]"}' *) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|forget} [flags]"}'
exit 2 ;; exit 2 ;;
esac esac
+8
View File
@@ -0,0 +1,8 @@
<!DOCTYPE html><html lang="fr"><head>
<title>Mon App</title>
<script type="module" crossorigin src="/assets/index-a1b2c3.js"></script>
<link rel="stylesheet" href="/assets/index-d4e5f6.css">
</head><body>
<div id="root"></div>
<script>window.__INITIAL_STATE__={"user":null,"routes":["/","/about","/contact"],"config":{"apiUrl":"https://api.example.com","features":["a","b","c"]}};</script>
</body></html>
+8
View File
@@ -0,0 +1,8 @@
<!DOCTYPE html><html lang="fr"><head>
<title>Lavage auto</title>
<meta name="description" content="Lavage auto à la main en Seine-et-Marne.">
<script type="application/ld+json">{"@context":"https://schema.org","@type":"LocalBusiness","name":"X"}</script>
</head><body>
<h1>Lavage auto à la main</h1>
<p>Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique. Lavage automobile à la main à Lagny-sur-Marne, detailing et protection céramique.</p>
</body></html>
+106
View File
@@ -0,0 +1,106 @@
#!/usr/bin/env python3
"""Is the content in the served HTML, or painted by JS? Stdlib only.
seo-analyzer records `RENDERING: SSR/SSG/SPA/hybrid` and then does nothing
with it. That is the gap this closes. On a client-rendered site `curl` returns
an empty shell, so every meta/H1/JSON-LD check reports "missing" and the audit
emits a page of false findings against a site that may be perfectly fine.
The verdict is taken from what the server actually sent — not from guessing at
package.json, where a React SPA and a Next.js SSR app look identical.
R2, not R1: this REPORTS blindness so the agent can refuse to score. It does
not render JS. No Playwright, no Chromium, no venv.
"""
import argparse, json, re
from html.parser import HTMLParser
import sitemap as sm # sibling: _fetch / _mock
# A shell can still carry a title + a couple of nav words. These thresholds
# separate "shell" from "page" on the two real sites measured 2026-07-17
# (server-rendered: 1 h1, thousands of body chars) and on a hydration stub.
MIN_TEXT = 400
MIN_H1 = 1
class _Doc(HTMLParser):
"""Collect body text and the tags an SEO audit reads. Script/style content
is NOT text: a 200 KB React bundle would otherwise look like a rich page."""
SKIP = ("script", "style", "noscript", "template", "svg")
def __init__(self):
super().__init__(convert_charrefs=True)
self.text, self.h1, self.jsonld, self.meta_desc = [], 0, 0, False
self._skip = 0
self._ld = False
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag in self.SKIP:
self._skip += 1
self._ld = tag == "script" and a.get("type") == "application/ld+json"
elif tag == "h1":
self.h1 += 1
elif tag == "meta" and a.get("name", "").lower() == "description":
self.meta_desc = bool((a.get("content") or "").strip())
def handle_endtag(self, tag):
if tag in self.SKIP and self._skip:
self._skip -= 1
self._ld = False
def handle_data(self, data):
if self._ld:
self.jsonld += 1
elif not self._skip:
s = data.strip()
if s:
self.text.append(s)
def _verdict(text_chars, h1, jsonld):
if text_chars >= MIN_TEXT and h1 >= MIN_H1:
return "server-rendered"
if text_chars < MIN_TEXT and h1 == 0 and jsonld == 0:
return "client-rendered"
return "partial" # shell + some SSR'd head, or thin page
def render_check(url):
raw = sm._mock("page.html")
if raw is None:
try:
raw = sm._fetch(url)
except Exception:
return {"status": "degraded", "reason": "fetch_failed"}
html = raw.decode("utf-8", "replace")
d = _Doc()
try:
d.feed(html)
except Exception:
pass # tolerate malformed markup
text = re.sub(r"\s+", " ", " ".join(d.text)).strip()
verdict = _verdict(len(text), d.h1, d.jsonld)
out = {"status": "ok", "source": "render_check", "verdict": verdict,
"body_text_chars": len(text), "h1_in_html": d.h1,
"jsonld_in_html": d.jsonld, "meta_description_in_html": d.meta_desc,
"html_bytes": len(raw)}
if verdict != "server-rendered":
out["warning"] = ("content is not in the served HTML — curl-based "
"on-page checks will report false 'missing' findings")
return out
def _cli():
try:
p = argparse.ArgumentParser()
p.add_argument("--url", required=True)
p.add_argument("--store", default=None) # accepted+ignored
args = p.parse_args()
print(json.dumps(render_check(args.url), indent=2))
except SystemExit as e:
if e.code not in (0, None):
print(json.dumps({"status": "error", "reason": "bad_usage"}))
raise
except Exception:
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
if __name__ == "__main__":
_cli()
+16
View File
@@ -135,6 +135,22 @@ has "billion-laughs refused" "$DTD" '"status": "degraded"'
has "dtd reason is distinct" "$DTD" 'unsafe_xml_dtd' has "dtd reason is distinct" "$DTD" 'unsafe_xml_dtd'
hasnt "dtd never parsed" "$DTD" '"count"' hasnt "dtd never parsed" "$DTD" '"count"'
echo "── render_check (R2) ──"
SPA="$(SEO_DATA_MOCK_DIR="$SD/fixtures-spa" python3 "$SD/render_check.py" \
--url https://spa.example/)"
has "spa → client-rendered" "$SPA" '"verdict": "client-rendered"'
has "spa has no h1 in html" "$SPA" '"h1_in_html": 0'
has "spa warns about false negs" "$SPA" 'false'
# the shell carries a fat window.__INITIAL_STATE__ script: script text is NOT
# page text, or a 200KB React bundle would read as a rich page
has "script text is not content" "$SPA" '"body_text_chars": 7'
SSR="$(SEO_DATA_MOCK_DIR="$SD/fixtures-ssr" python3 "$SD/render_check.py" \
--url https://ssr.example/)"
has "ssr → server-rendered" "$SSR" '"verdict": "server-rendered"'
has "ssr counts jsonld" "$SSR" '"jsonld_in_html": 1'
has "ssr sees meta description" "$SSR" '"meta_description_in_html": true'
hasnt "ssr emits no warning" "$SSR" 'warning'
echo "── linkgraph ──" echo "── linkgraph ──"
LG="$(SEO_DATA_MOCK_DIR="$SD/fixtures-linkgraph" python3 "$SD/linkgraph.py" \ LG="$(SEO_DATA_MOCK_DIR="$SD/fixtures-linkgraph" python3 "$SD/linkgraph.py" \
--url https://ex.com/sitemap.xml)" --url https://ex.com/sitemap.xml)"