From 4818c6116f6d1b327bd60c4296515d91f1dc34aa Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 13:29:14 +0200 Subject: [PATCH] =?UTF-8?q?feat(seo-data):=20I7=20=E2=80=94=20compute=20th?= =?UTF-8?q?e=20score=20instead=20of=20feeling=20it?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit /harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3, Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over identical code could disagree. That is a credibility problem on its own, and /client-handover gates on 17/20 — a wobbling number makes the gate arbitrary. H2 sharpened it: now that drift reports what actually changed, a score moving on its own is visibly noise. The split is the whole point. WHICH findings exist and how severe each is stays the LLM's judgement — irreducible, and I am not pretending otherwise. The arithmetic stops being judgement: same findings in, same score out. Same principle as grouping cannibalisation rows in the engine rather than handing a model 1000 rows to add up. Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary instead of two. Two things it makes real that were prose: - **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable off-page) both mandate excluding an axis and renormalising the rest. Both left that arithmetic to the model. Now the engine does it and refuses to let N/A behave like a zero — verified: all-20 axes with two N/A still yields global 20.0, not a dragged-down mean. - **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a single page de-escalates). A defect on 1 of 12 pages is not the defect on 12 of 12, and flattening the two is part of what made the old numbers move. Malformed input is an error, never a silently wrong number — unlike the fetch verbs, a degrade here would mean bad input, not a network fact. Unknown severity and unknown profile both rejected, tested. Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8; critique+haute = 77 → 15.4), identical global across repeated runs, weights renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail; full suite green; shellcheck + py_compile clean. --- agents/seo-analyzer.md | 35 +++++++++++ lib/seo-data/README.md | 22 +++++++ lib/seo-data/fetch.sh | 4 +- lib/seo-data/score.py | 113 ++++++++++++++++++++++++++++++++++ lib/seo-data/seo-data.test.sh | 29 +++++++++ 5 files changed, 202 insertions(+), 1 deletion(-) create mode 100644 lib/seo-data/score.py diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 6ec050d..82cb309 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -898,6 +898,41 @@ FIX: AUTO () | USER () | Competitive position | 5% | 10% | | | Legal compliance | 10% | 5% | | +**Compute the scores, do not feel them (I7).** Emit your findings, then let +the engine do the arithmetic: + +```bash +bash ~/.claude/lib/seo-data/fetch.sh score --findings /tmp/seo-findings.json +``` + +```json +{"depth":"FULL","profile":"local", + "axes":{"technical":{"findings":[{"severity":"haute","affected":9,"sampled":12}]}, + "on-page":{"status":"na","reason":"client-rendered (R2)"}, + "off-page":{"status":"na","reason":"backlinks unauditable (I1)"}}} +``` + +`profile`: `local` (B2C) | `national` (SaaS/national/content). Severities are +`critique|haute|moyenne|basse` — `/harden`'s scale (-15/-8/-3/-1, clamp, +then /5 into /20), so the whole skill family speaks one vocabulary. + +**The split matters.** WHICH findings exist and how severe each is stays your +judgement — irreducible. The addition is not: same findings in, same score +out. Until now every axis was felt, so two runs over identical code could +disagree, and `/client-handover` gates on 17/20. + +- `affected`/`sampled` (optional) shift severity ONE step: ≥50% of the sample + escalates, a single page de-escalates. A defect on 1 of 12 pages is not the + defect on 12 of 12; pretending so is what made the old numbers wobble. +- `status: "na"` → the axis is EXCLUDED and the remaining weights are + renormalised for you. This is the R2 rule (client-rendered on-page) and the + I1 rule (unauditable off-page), finally computed instead of done by hand. + **N/A is not a zero** and the engine will not let it behave like one. +- `status: "error"` → malformed findings. Fix them; never fall back to + eyeballing a number. +- Run it twice on the same file before publishing. If the output moved, your + findings moved, and that is the thing to explain. + **Technical axis note:** CWV scored on CrUX field data (75th percentile, real users, from STEP 4) when available; otherwise lab PageSpeed Lighthouse run. diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index c4e11e9..747a244 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -193,6 +193,28 @@ fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500] • Mock is pages.json ({url: html}), not a single page.html: one fixture cannot express a graph — every node would carry identical links. +fetch.sh score --findings + → {"status":"ok","axes":{"technical":{"score_20":17.8,"weight":0.2, + "weight_renormalised":0.2857,"findings":2}}, + "na":["off-page","on-page"],"weights_renormalised":true,"global_20":17.6} + → {"status":"error","reason":"unknown severity: 'bogus'"|"bad_findings_json"} + + I7. /harden has a real scale (SKILL.md:435: -15/-8/-3/-1, clamp [0,100]); + /seo had none, so every axis was FELT and two runs over identical code could + disagree — while /client-handover gates on 17/20. Same scale here, /5 into + /20, one vocabulary across the family. + • The split: WHICH findings exist and how severe each is stays the LLM's + judgement. The addition is not. Same findings in, same score out. + • affected/sampled shift severity ONE step: >=50% of the sample escalates, + a single page de-escalates. A defect on 1 of 12 pages is not the defect + on 12 of 12. + • status:"na" → axis EXCLUDED, remaining weights renormalised. This is + R2's rule (client-rendered on-page) and I1's (unauditable off-page), + computed rather than done by hand. N/A is not a zero, and the engine + will not let it act like one. + • Malformed input is an error, never a silently wrong number — unlike the + fetch verbs, a degrade here would mean bad input, not a network fact. + fetch.sh drift --url https://ex.com/sitemap.xml [--max 500] → {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"} → {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…], diff --git a/lib/seo-data/fetch.sh b/lib/seo-data/fetch.sh index 546f4eb..e2291be 100644 --- a/lib/seo-data/fetch.sh +++ b/lib/seo-data/fetch.sh @@ -32,6 +32,8 @@ case "$cmd" in # No auth, no Google: stdlib-only, runs even without the venv. sitemap) exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;; + score) + exec "$PY" "$HERE/score.py" --store "$STORE" "$@" ;; drift) exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;; rendercheck) @@ -50,6 +52,6 @@ case "$cmd" in fi echo '{"status":"error","reason":"usage: fetch.sh forget {--label