Files
Bastien Chanot 20d3082542 feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright
Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
2026-07-17 12:44:43 +02:00

107 lines
4.0 KiB
Python

#!/usr/bin/env python3
"""Is the content in the served HTML, or painted by JS? Stdlib only.
seo-analyzer records `RENDERING: SSR/SSG/SPA/hybrid` and then does nothing
with it. That is the gap this closes. On a client-rendered site `curl` returns
an empty shell, so every meta/H1/JSON-LD check reports "missing" and the audit
emits a page of false findings against a site that may be perfectly fine.
The verdict is taken from what the server actually sent — not from guessing at
package.json, where a React SPA and a Next.js SSR app look identical.
R2, not R1: this REPORTS blindness so the agent can refuse to score. It does
not render JS. No Playwright, no Chromium, no venv.
"""
import argparse, json, re
from html.parser import HTMLParser
import sitemap as sm # sibling: _fetch / _mock
# A shell can still carry a title + a couple of nav words. These thresholds
# separate "shell" from "page" on the two real sites measured 2026-07-17
# (server-rendered: 1 h1, thousands of body chars) and on a hydration stub.
MIN_TEXT = 400
MIN_H1 = 1
class _Doc(HTMLParser):
"""Collect body text and the tags an SEO audit reads. Script/style content
is NOT text: a 200 KB React bundle would otherwise look like a rich page."""
SKIP = ("script", "style", "noscript", "template", "svg")
def __init__(self):
super().__init__(convert_charrefs=True)
self.text, self.h1, self.jsonld, self.meta_desc = [], 0, 0, False
self._skip = 0
self._ld = False
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag in self.SKIP:
self._skip += 1
self._ld = tag == "script" and a.get("type") == "application/ld+json"
elif tag == "h1":
self.h1 += 1
elif tag == "meta" and a.get("name", "").lower() == "description":
self.meta_desc = bool((a.get("content") or "").strip())
def handle_endtag(self, tag):
if tag in self.SKIP and self._skip:
self._skip -= 1
self._ld = False
def handle_data(self, data):
if self._ld:
self.jsonld += 1
elif not self._skip:
s = data.strip()
if s:
self.text.append(s)
def _verdict(text_chars, h1, jsonld):
if text_chars >= MIN_TEXT and h1 >= MIN_H1:
return "server-rendered"
if text_chars < MIN_TEXT and h1 == 0 and jsonld == 0:
return "client-rendered"
return "partial" # shell + some SSR'd head, or thin page
def render_check(url):
raw = sm._mock("page.html")
if raw is None:
try:
raw = sm._fetch(url)
except Exception:
return {"status": "degraded", "reason": "fetch_failed"}
html = raw.decode("utf-8", "replace")
d = _Doc()
try:
d.feed(html)
except Exception:
pass # tolerate malformed markup
text = re.sub(r"\s+", " ", " ".join(d.text)).strip()
verdict = _verdict(len(text), d.h1, d.jsonld)
out = {"status": "ok", "source": "render_check", "verdict": verdict,
"body_text_chars": len(text), "h1_in_html": d.h1,
"jsonld_in_html": d.jsonld, "meta_description_in_html": d.meta_desc,
"html_bytes": len(raw)}
if verdict != "server-rendered":
out["warning"] = ("content is not in the served HTML — curl-based "
"on-page checks will report false 'missing' findings")
return out
def _cli():
try:
p = argparse.ArgumentParser()
p.add_argument("--url", required=True)
p.add_argument("--store", default=None) # accepted+ignored
args = p.parse_args()
print(json.dumps(render_check(args.url), indent=2))
except SystemExit as e:
if e.code not in (0, None):
print(json.dumps({"status": "error", "reason": "bad_usage"}))
raise
except Exception:
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
if __name__ == "__main__":
_cli()