feat(seo-data): H2 — drift baseline; regressions vs changes, not prose

seo-analyzer.md:1365 keeps history as "date + score + key changes" — prose the
LLM writes about its own previous prose. Lossy, unreproducible, and
machine-uncomparable, so "the redesign silently dropped 40 canonicals" is
invisible unless someone happens to notice.

drift snapshots title/description/canonical/robots/h1_count/jsonld_types per
URL and diffs them. Stdlib only, no auth.

The classification IS the feature: LOSING a signal is a regression, CHANGING
one is a change that may well be intended. The engine says which kind; the
agent judges. A reworded title is not an alert; an evaporated canonical is.

Runs over the WHOLE sitemap, never a sample — caught while designing: a drift
computed over a sample that changes between runs compares nothing.

NOT rank tracking. That is the common misread of this same feature elsewhere;
positions come from GSC `queries`. This is on-page regression detection.

Also caught in my own draft before testing: _capture reused
sm._mock("page.html"), the exact single-fixture flaw I had already fixed in
linkgraph — one fixture cannot express a multi-page snapshot, every URL would
read identical. Now pages.json, same convention.

Proved on a planted failure rather than a happy path — two clean sites would
look identical to a detector that always returns []:
  v1 -> v2: canonical lost on /a, h1 + jsonld lost on /, title reworded,
  /gone removed, /neuve added
  → 3 regressions, 1 change, gone/new both detected, title correctly NOT a
    regression.

Store is ~/.claude/seo-data/drift/<host>.json, 0700, written via os.replace so
a crash never leaves a half-written baseline; a corrupt store degrades to
"first run" instead of killing the audit.

Verified: seo-data 144 -> 155 pass, 0 fail; full suite green.
This commit is contained in:
Bastien Chanot
2026-07-17 13:25:45 +02:00
parent d6b8edc8ea
commit f69cfc5cb4
8 changed files with 256 additions and 1 deletions
+21
View File
@@ -193,6 +193,27 @@ fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500]
• Mock is pages.json ({url: html}), not a single page.html: one fixture • Mock is pages.json ({url: html}), not a single page.html: one fixture
cannot express a graph — every node would carry identical links. cannot express a graph — every node would carry identical links.
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
"regressions":[{"url":…,"field":"canonical","was":"…","now":null}],
"changes":[{"url":…,"field":"title","was":"…","now":"…"}]}
On-page drift between audits. seo-analyzer.md:1365 keeps only "date + score
+ key changes" as PROSE the LLM writes about its own previous prose: lossy,
unreproducible, machine-uncomparable. So "the redesign silently dropped 40
canonicals" stays invisible. This snapshots title/description/canonical/
robots/h1_count/jsonld_types per URL and diffs them.
• NOT rank tracking (the common misread of this feature elsewhere).
Positions come from GSC `queries`. This is regression detection.
• Runs over the WHOLE sitemap, never a sample: a drift over a sample that
changes between runs compares nothing.
• LOSING a signal = regression. CHANGING one = change, possibly intended —
the agent judges that, the engine only says which kind it is.
• Store: ~/.claude/seo-data/drift/<host>.json, 0700, written via
os.replace — never a half-written baseline. Corrupt store → treated as
a first run rather than crashing the audit.
fetch.sh forget --label client-a fetch.sh forget --label client-a
→ {"status":"ok","removed":true|false} # false = label wasn't in the store → {"status":"ok","removed":true|false} # false = label wasn't in the store
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env python3
"""On-page drift between audits. Stdlib only.
seo-analyzer.md:1365 says "on re-run, move current content to Historique
(summary: date + score + key changes)". That is prose the LLM writes about its
own previous prose: lossy, unreproducible, and machine-uncomparable. So "the
redesign silently dropped 40 canonicals" is invisible unless someone happens
to notice.
This snapshots the machine-readable signals per URL and diffs them.
NOT rank tracking — a common misread of the same feature elsewhere. Positions
come from GSC (`queries`). This is on-page regression detection: what the site
said last time vs now.
Runs over the WHOLE sitemap, never a sample: a drift over a sample that
changes between runs compares nothing.
"""
import argparse, json, os, re, time
from html.parser import HTMLParser
import sitemap as sm
STORE_DIR = os.path.expanduser("~/.claude/seo-data/drift")
MAX_PAGES = 500
# Losing a signal is a regression. Changing one may be intentional — the agent
# judges that, we only report which kind it is.
TRACKED = ("title", "description", "canonical", "robots", "h1_count", "jsonld_types")
class _Signals(HTMLParser):
def __init__(self):
super().__init__(convert_charrefs=True)
self.title, self.description, self.canonical, self.robots = None, None, None, None
self.h1_count, self.jsonld_types = 0, []
self._in_title, self._in_ld = False, False
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag == "title":
self._in_title = True
elif tag == "h1":
self.h1_count += 1
elif tag == "meta":
n = (a.get("name") or "").lower()
if n == "description":
self.description = (a.get("content") or "").strip() or None
elif n == "robots":
self.robots = (a.get("content") or "").strip() or None
elif tag == "link" and "canonical" in (a.get("rel") or "").lower():
self.canonical = (a.get("href") or "").strip() or None
elif tag == "script" and a.get("type") == "application/ld+json":
self._in_ld = True
def handle_endtag(self, tag):
if tag == "title":
self._in_title = False
elif tag == "script":
self._in_ld = False
def handle_data(self, data):
if self._in_title and data.strip():
self.title = re.sub(r"\s+", " ", data.strip())
elif self._in_ld:
self.jsonld_types.extend(re.findall(r'"@type"\s*:\s*"([^"]+)"', data))
def _signals(html):
p = _Signals()
try:
p.feed(html)
except Exception:
pass
return {"title": p.title, "description": p.description,
"canonical": p.canonical, "robots": p.robots,
"h1_count": p.h1_count, "jsonld_types": sorted(set(p.jsonld_types))}
def _mock_pages():
"""{url: html}, same convention as linkgraph: a single page.html fixture
cannot express a multi-page snapshot — every URL would look identical."""
raw = sm._mock("pages.json")
return json.loads(raw.decode("utf-8")) if raw else None
def _capture(urls):
pages = _mock_pages()
snap, failed = {}, 0
for u in urls:
if pages is not None:
html = pages.get(u)
if html is None:
failed += 1
continue
else:
try:
html = sm._fetch(u).decode("utf-8", "replace")
except Exception:
failed += 1
continue
snap[u] = _signals(html)
return snap, failed
def _store_path(sitemap_url):
from urllib.parse import urlparse
host = urlparse(sitemap_url).netloc.lower()
safe = re.sub(r"[^a-z0-9.-]", "_", host) or "unknown"
return os.path.join(STORE_DIR, safe + ".json")
def _load(path):
if not os.path.exists(path):
return None
try:
with open(path, encoding="utf-8") as f:
return json.load(f)
except Exception:
return None # corrupt store -> treat as first run
def _save(path, snap, stamp):
os.makedirs(os.path.dirname(path), mode=0o700, exist_ok=True)
tmp = path + ".tmp"
with open(tmp, "w", encoding="utf-8") as f:
json.dump({"captured": stamp, "pages": snap}, f)
os.replace(tmp, path) # atomic: never a half-written baseline
def _classify(old, new):
"""LOST a signal = regression. Changed it = change. Only the first is
unambiguous; the agent judges the rest."""
regressions, changes = [], []
for f in TRACKED:
o, n = old.get(f), new.get(f)
if o == n:
continue
row = {"field": f, "was": o, "now": n}
# Covers every tracked field uniformly: "Titre" -> None, 1 -> 0,
# ["Article"] -> []. Had the value, lost the value.
(regressions if (o and not n) else changes).append(row)
return regressions, changes
def drift(sitemap_url, max_pages=MAX_PAGES):
sm_res = sm.sitemap(sitemap_url)
if sm_res.get("status") != "ok":
return sm_res
urls = sm_res["urls"][:max_pages]
snap, failed = _capture(urls)
if not snap:
return {"status": "degraded", "reason": "no_pages_fetched"}
stamp = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
path = _store_path(sitemap_url)
prev = _load(path)
_save(path, snap, stamp)
if prev is None:
return {"status": "ok", "baseline": True, "captured": stamp,
"pages": len(snap), "pages_failed": failed, "store": path}
old = prev.get("pages", {})
regressions, changes = [], []
for u, new in snap.items():
if u not in old:
continue
r, c = _classify(old[u], new)
for row in r:
regressions.append(dict(row, url=u))
for row in c:
changes.append(dict(row, url=u))
return {"status": "ok", "baseline": False,
"since": prev.get("captured"), "captured": stamp,
"pages": len(snap), "pages_failed": failed,
"gone": sorted(set(old) - set(snap)),
"new": sorted(set(snap) - set(old)),
"regressions": regressions, "changes": changes, "store": path}
def _cli():
try:
p = argparse.ArgumentParser()
p.add_argument("--url", required=True, help="sitemap URL")
p.add_argument("--max", type=int, default=MAX_PAGES)
p.add_argument("--store", default=None) # accepted+ignored
args = p.parse_args()
print(json.dumps(drift(args.url, args.max), indent=2))
except SystemExit as e:
if e.code not in (0, None):
print(json.dumps({"status": "error", "reason": "bad_usage"}))
raise
except Exception:
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
if __name__ == "__main__":
_cli()
+3 -1
View File
@@ -32,6 +32,8 @@ case "$cmd" in
# No auth, no Google: stdlib-only, runs even without the venv. # No auth, no Google: stdlib-only, runs even without the venv.
sitemap) sitemap)
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;; exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
drift)
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
rendercheck) rendercheck)
exec "$PY" "$HERE/render_check.py" --store "$STORE" "$@" ;; exec "$PY" "$HERE/render_check.py" --store "$STORE" "$@" ;;
linkgraph) linkgraph)
@@ -48,6 +50,6 @@ case "$cmd" in
fi fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}' echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;; exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|forget} [flags]"}' *) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|forget} [flags]"}'
exit 2 ;; exit 2 ;;
esac esac
@@ -0,0 +1,5 @@
{
"https://ex.com/": "<html><head><title>Accueil</title><meta name='description' content='desc'><link rel='canonical' href='https://ex.com/'><script type='application/ld+json'>{\"@type\":\"LocalBusiness\"}</script></head><body><h1>Accueil</h1></body></html>",
"https://ex.com/a": "<html><head><title>Page A</title><link rel='canonical' href='https://ex.com/a'></head><body><h1>A</h1></body></html>",
"https://ex.com/gone": "<html><head><title>Bientot supprimee</title></head><body><h1>G</h1></body></html>"
}
@@ -0,0 +1,6 @@
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url><loc>https://ex.com/</loc></url>
<url><loc>https://ex.com/a</loc></url>
<url><loc>https://ex.com/gone</loc></url>
</urlset>
@@ -0,0 +1,5 @@
{
"https://ex.com/": "<html><head><title>Accueil refondue</title><meta name='description' content='desc'><link rel='canonical' href='https://ex.com/'></head><body><p>plus de h1, plus de jsonld</p></body></html>",
"https://ex.com/a": "<html><head><title>Page A</title></head><body><h1>A</h1></body></html>",
"https://ex.com/neuve": "<html><head><title>Neuve</title></head><body><h1>N</h1></body></html>"
}
@@ -0,0 +1,6 @@
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url><loc>https://ex.com/</loc></url>
<url><loc>https://ex.com/a</loc></url>
<url><loc>https://ex.com/neuve</loc></url>
</urlset>
+26
View File
@@ -178,6 +178,32 @@ has "cap is reported" "$CAP" '"capped": true'
has "capped withholds orphans" "$CAP" '"orphans_withheld": true' has "capped withholds orphans" "$CAP" '"orphans_withheld": true'
hasnt "capped emits no orphans" "$CAP" '"orphans":' hasnt "capped emits no orphans" "$CAP" '"orphans":'
echo "── drift (H2) ──"
DH="$(mktemp -d)"
D1="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v1" python3 "$SD/drift.py" \
--url https://ex.com/sitemap.xml)"
has "first run is a baseline" "$D1" '"baseline": true'
has "baseline captures pages" "$D1" '"pages": 3'
hasnt "baseline diffs nothing" "$D1" '"regressions"'
# v2: canonical lost on /a, h1+jsonld lost on /, title reworded, /gone removed,
# /neuve added. Losses are regressions; a reworded title is not.
D2="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v2" python3 "$SD/drift.py" \
--url https://ex.com/sitemap.xml)"
has "second run diffs" "$D2" '"baseline": false'
has "detects removed url" "$D2" '"https://ex.com/gone"'
has "detects added url" "$D2" '"https://ex.com/neuve"'
has "lost canonical = regression" "$D2" '"canonical"'
has "lost h1 = regression" "$D2" '"h1_count"'
has "lost jsonld = regression" "$D2" '"jsonld_types"'
# the classification IS the feature: losing a signal != changing one
NREG="$(printf '%s' "$D2" | python3 -c 'import sys,json; print(len(json.load(sys.stdin)["regressions"]))')"
NCHG="$(printf '%s' "$D2" | python3 -c 'import sys,json; print(len(json.load(sys.stdin)["changes"]))')"
[ "$NREG" = "3" ] && ok "3 losses classed as regressions" \
|| no "3 losses classed as regressions" "got $NREG"
[ "$NCHG" = "1" ] && ok "reworded title is a change, not a regression" \
|| no "reworded title is a change, not a regression" "got $NCHG"
rm -rf "$DH"
echo "── fetch.sh ──" echo "── fetch.sh ──"
FETCH="$SD/fetch.sh" FETCH="$SD/fetch.sh"
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env — # SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —