feat(seo-data): I7 — compute the score instead of feeling it

/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.

The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.

Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.

Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
  off-page) both mandate excluding an axis and renormalising the rest. Both
  left that arithmetic to the model. Now the engine does it and refuses to let
  N/A behave like a zero — verified: all-20 axes with two N/A still yields
  global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
  single page de-escalates). A defect on 1 of 12 pages is not the defect on
  12 of 12, and flattening the two is part of what made the old numbers move.

Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.

Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 13:29:14 +02:00
parent f69cfc5cb4
commit 4818c6116f
5 changed files with 202 additions and 1 deletions
+35
View File
@@ -898,6 +898,41 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
| Competitive position | 5% | 10% | |
| Legal compliance | 10% | 5% | |
**Compute the scores, do not feel them (I7).** Emit your findings, then let
the engine do the arithmetic:
```bash
bash ~/.claude/lib/seo-data/fetch.sh score --findings /tmp/seo-findings.json
```
```json
{"depth":"FULL","profile":"local",
"axes":{"technical":{"findings":[{"severity":"haute","affected":9,"sampled":12}]},
"on-page":{"status":"na","reason":"client-rendered (R2)"},
"off-page":{"status":"na","reason":"backlinks unauditable (I1)"}}}
```
`profile`: `local` (B2C) | `national` (SaaS/national/content). Severities are
`critique|haute|moyenne|basse` — `/harden`'s scale (-15/-8/-3/-1, clamp,
then /5 into /20), so the whole skill family speaks one vocabulary.
**The split matters.** WHICH findings exist and how severe each is stays your
judgement — irreducible. The addition is not: same findings in, same score
out. Until now every axis was felt, so two runs over identical code could
disagree, and `/client-handover` gates on 17/20.
- `affected`/`sampled` (optional) shift severity ONE step: ≥50% of the sample
escalates, a single page de-escalates. A defect on 1 of 12 pages is not the
defect on 12 of 12; pretending so is what made the old numbers wobble.
- `status: "na"` → the axis is EXCLUDED and the remaining weights are
renormalised for you. This is the R2 rule (client-rendered on-page) and the
I1 rule (unauditable off-page), finally computed instead of done by hand.
**N/A is not a zero** and the engine will not let it behave like one.
- `status: "error"` → malformed findings. Fix them; never fall back to
eyeballing a number.
- Run it twice on the same file before publishing. If the output moved, your
findings moved, and that is the thing to explain.
**Technical axis note:** CWV scored on CrUX field data (75th percentile,
real users, from STEP 4) when available; otherwise lab PageSpeed
Lighthouse run.
+22
View File
@@ -193,6 +193,28 @@ fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500]
• Mock is pages.json ({url: html}), not a single page.html: one fixture
cannot express a graph — every node would carry identical links.
fetch.sh score --findings <path.json | ->
→ {"status":"ok","axes":{"technical":{"score_20":17.8,"weight":0.2,
"weight_renormalised":0.2857,"findings":2}},
"na":["off-page","on-page"],"weights_renormalised":true,"global_20":17.6}
→ {"status":"error","reason":"unknown severity: 'bogus'"|"bad_findings_json"}
I7. /harden has a real scale (SKILL.md:435: -15/-8/-3/-1, clamp [0,100]);
/seo had none, so every axis was FELT and two runs over identical code could
disagree — while /client-handover gates on 17/20. Same scale here, /5 into
/20, one vocabulary across the family.
• The split: WHICH findings exist and how severe each is stays the LLM's
judgement. The addition is not. Same findings in, same score out.
• affected/sampled shift severity ONE step: >=50% of the sample escalates,
a single page de-escalates. A defect on 1 of 12 pages is not the defect
on 12 of 12.
• status:"na" → axis EXCLUDED, remaining weights renormalised. This is
R2's rule (client-rendered on-page) and I1's (unauditable off-page),
computed rather than done by hand. N/A is not a zero, and the engine
will not let it act like one.
• Malformed input is an error, never a silently wrong number — unlike the
fetch verbs, a degrade here would mean bad input, not a network fact.
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
+3 -1
View File
@@ -32,6 +32,8 @@ case "$cmd" in
# No auth, no Google: stdlib-only, runs even without the venv.
sitemap)
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
score)
exec "$PY" "$HERE/score.py" --store "$STORE" "$@" ;;
drift)
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
rendercheck)
@@ -50,6 +52,6 @@ case "$cmd" in
fi
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
exit 2 ;;
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|forget} [flags]"}'
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|forget} [flags]"}'
exit 2 ;;
esac
+113
View File
@@ -0,0 +1,113 @@
#!/usr/bin/env python3
"""Deterministic /20 scoring from a findings list. Stdlib only.
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo has none: every axis is felt, not computed, so
two runs over identical code can produce different scores. That is a
credibility problem on its own, and /client-handover gates on 17/20 — a
wobbling number makes the gate arbitrary. H2 sharpens it further: now that
drift reports what actually changed, a score moving on its own is visibly
noise.
The split is the point. The LLM keeps the irreducible judgement — WHICH
findings exist and how severe each is. The arithmetic stops being judgement:
same findings in, same score out. Same principle as grouping cannibalisation
rows in the engine rather than asking a model to add up 1000 of them.
Scale is /harden's, /5 into /20, so the whole skill family speaks one
vocabulary.
"""
import argparse, json, sys
PENALTY = {"critique": 15, "haute": 8, "moyenne": 3, "basse": 1}
# STEP 9 weights. FULL = 7 axes, LOCAL = 4 (off-page/social/competitive are
# not audited at that depth).
WEIGHTS = {
("FULL", "local"): {"technical": .20, "on-page": .20, "seo-local": .25,
"off-page": .10, "social": .10, "competitive": .05,
"legal": .10},
("FULL", "national"): {"technical": .30, "on-page": .30, "seo-local": .05,
"off-page": .15, "social": .05, "competitive": .10,
"legal": .05},
("LOCAL", "local"): {"technical": .25, "on-page": .35, "seo-local": .20,
"legal": .20},
("LOCAL", "national"):{"technical": .35, "on-page": .45, "seo-local": .05,
"legal": .15},
}
def _axis_score(findings):
"""100 - Σ penalties, clamped, then /5 → /20. Prevalence shifts severity
ONE step, never invents one: a finding on 1 of 12 sampled pages is not the
same defect as one on 12 of 12, and pretending otherwise is what made the
old scores unreproducible."""
total = 0
for f in findings:
sev = str(f.get("severity", "")).lower()
if sev not in PENALTY:
raise ValueError("unknown severity: %r" % f.get("severity"))
order = ["basse", "moyenne", "haute", "critique"]
i = order.index(sev)
aff, samp = f.get("affected"), f.get("sampled")
if isinstance(aff, int) and isinstance(samp, int) and samp > 0:
ratio = aff / samp
if ratio >= 0.5:
i = min(i + 1, len(order) - 1) # widespread → escalate
elif aff <= 1:
i = max(i - 1, 0) # isolated → de-escalate
total += PENALTY[order[i]]
return round(max(0, 100 - total) / 5.0, 1)
def score(payload):
depth = str(payload.get("depth", "FULL")).upper()
profile = str(payload.get("profile", "local")).lower()
key = (depth, profile)
if key not in WEIGHTS:
return {"status": "error", "reason": "unknown depth/profile: %s/%s"
% (depth, profile)}
weights, axes_in = WEIGHTS[key], payload.get("axes", {})
scored, na = {}, []
for axis, w in weights.items():
a = axes_in.get(axis)
if a is None or str(a.get("status", "")).lower() == "na":
na.append(axis) # N/A is not a zero
continue
try:
s = _axis_score(a.get("findings", []))
except ValueError as e:
return {"status": "error", "reason": str(e)}
scored[axis] = {"score_20": s, "weight": w,
"findings": len(a.get("findings", []))}
if not scored:
return {"status": "degraded", "reason": "no_axis_scored"}
# Renormalise over what was actually measured. R2 mandates this for a
# client-rendered on-page axis and left it to the model to do by hand.
live = sum(v["weight"] for v in scored.values())
for v in scored.values():
v["weight_renormalised"] = round(v["weight"] / live, 4)
glob = sum(v["score_20"] * v["weight"] / live for v in scored.values())
return {"status": "ok", "source": "score", "depth": depth,
"profile": profile, "axes": scored, "na": sorted(na),
"weights_renormalised": round(live, 4) != 1.0,
"global_20": round(glob, 1)}
def _cli():
try:
p = argparse.ArgumentParser()
p.add_argument("--findings", default="-", help="JSON path, or - for stdin")
p.add_argument("--store", default=None) # accepted+ignored
args = p.parse_args()
raw = sys.stdin.read() if args.findings == "-" else \
open(args.findings, encoding="utf-8").read()
print(json.dumps(score(json.loads(raw)), indent=2))
except SystemExit as e:
if e.code not in (0, None):
print(json.dumps({"status": "error", "reason": "bad_usage"}))
raise
except Exception:
# Unlike the fetch verbs this is pure arithmetic: a degrade here means
# malformed input, never a network fact.
print(json.dumps({"status": "error", "reason": "bad_findings_json"}))
if __name__ == "__main__":
_cli()
+29
View File
@@ -178,6 +178,35 @@ has "cap is reported" "$CAP" '"capped": true'
has "capped withholds orphans" "$CAP" '"orphans_withheld": true'
hasnt "capped emits no orphans" "$CAP" '"orphans":'
echo "── score (I7) ──"
sc() { printf '%s' "$1" | python3 "$SD/score.py" --findings -; }
# technical: haute(-8) + moyenne(-3) = 100-11 = 89 → 17.8
B='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"haute"},{"severity":"moyenne"}]},"seo-local":{"findings":[]},"off-page":{"findings":[]},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]},"on-page":{"findings":[]}}}'
R="$(sc "$B")"
has "harden scale, /5 into /20" "$R" '"score_20": 17.8'
has "no findings = 20" "$R" '"score_20": 20.0'
has "nothing renormalised" "$R" '"weights_renormalised": false'
# THE point of I7: same findings in, same score out
A1="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
A2="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
[ "$A1" = "$A2" ] && ok "score is reproducible" || no "score is reproducible" "$A1 vs $A2"
# N/A is not a zero, and R2 mandated renormalising by hand — now computed
NA='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[]},"on-page":{"status":"na"},"seo-local":{"findings":[]},"off-page":{"status":"na"},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]}}}'
RN="$(sc "$NA")"
has "na axes listed" "$RN" '"on-page"'
has "renormalisation flagged" "$RN" '"weights_renormalised": true'
# all axes 20 → global must stay 20: N/A must not drag the mean down
has "na is not a zero" "$RN" '"global_20": 20.0'
# prevalence shifts severity ONE step, both ways
WIDE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":10,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
ONE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":1,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
has "widespread escalates (-8)" "$(sc "$WIDE")" '"score_20": 18.4'
has "isolated de-escalates (-1)" "$(sc "$ONE")" '"score_20": 19.8'
# malformed input is an error, never a silently wrong number
has "unknown severity rejected" "$(sc '{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"bogus"}]}}}')" '"status": "error"'
has "unknown profile rejected" "$(sc '{"depth":"FULL","profile":"martian","axes":{}}')" '"status": "error"'
has "garbage json is an error" "$(sc 'not json')" '"status": "error"'
echo "── drift (H2) ──"
DH="$(mktemp -d)"
D1="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v1" python3 "$SD/drift.py" \