feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3, Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over identical code could disagree. That is a credibility problem on its own, and /client-handover gates on 17/20 — a wobbling number makes the gate arbitrary. H2 sharpened it: now that drift reports what actually changed, a score moving on its own is visibly noise. The split is the whole point. WHICH findings exist and how severe each is stays the LLM's judgement — irreducible, and I am not pretending otherwise. The arithmetic stops being judgement: same findings in, same score out. Same principle as grouping cannibalisation rows in the engine rather than handing a model 1000 rows to add up. Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary instead of two. Two things it makes real that were prose: - **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable off-page) both mandate excluding an axis and renormalising the rest. Both left that arithmetic to the model. Now the engine does it and refuses to let N/A behave like a zero — verified: all-20 axes with two N/A still yields global 20.0, not a dragged-down mean. - **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a single page de-escalates). A defect on 1 of 12 pages is not the defect on 12 of 12, and flattening the two is part of what made the old numbers move. Malformed input is an error, never a silently wrong number — unlike the fetch verbs, a degrade here would mean bad input, not a network fact. Unknown severity and unknown profile both rejected, tested. Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8; critique+haute = 77 → 15.4), identical global across repeated runs, weights renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail; full suite green; shellcheck + py_compile clean.
This commit is contained in:
@@ -898,6 +898,41 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
|||||||
| Competitive position | 5% | 10% | |
|
| Competitive position | 5% | 10% | |
|
||||||
| Legal compliance | 10% | 5% | |
|
| Legal compliance | 10% | 5% | |
|
||||||
|
|
||||||
|
**Compute the scores, do not feel them (I7).** Emit your findings, then let
|
||||||
|
the engine do the arithmetic:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash ~/.claude/lib/seo-data/fetch.sh score --findings /tmp/seo-findings.json
|
||||||
|
```
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"depth":"FULL","profile":"local",
|
||||||
|
"axes":{"technical":{"findings":[{"severity":"haute","affected":9,"sampled":12}]},
|
||||||
|
"on-page":{"status":"na","reason":"client-rendered (R2)"},
|
||||||
|
"off-page":{"status":"na","reason":"backlinks unauditable (I1)"}}}
|
||||||
|
```
|
||||||
|
|
||||||
|
`profile`: `local` (B2C) | `national` (SaaS/national/content). Severities are
|
||||||
|
`critique|haute|moyenne|basse` — `/harden`'s scale (-15/-8/-3/-1, clamp,
|
||||||
|
then /5 into /20), so the whole skill family speaks one vocabulary.
|
||||||
|
|
||||||
|
**The split matters.** WHICH findings exist and how severe each is stays your
|
||||||
|
judgement — irreducible. The addition is not: same findings in, same score
|
||||||
|
out. Until now every axis was felt, so two runs over identical code could
|
||||||
|
disagree, and `/client-handover` gates on 17/20.
|
||||||
|
|
||||||
|
- `affected`/`sampled` (optional) shift severity ONE step: ≥50% of the sample
|
||||||
|
escalates, a single page de-escalates. A defect on 1 of 12 pages is not the
|
||||||
|
defect on 12 of 12; pretending so is what made the old numbers wobble.
|
||||||
|
- `status: "na"` → the axis is EXCLUDED and the remaining weights are
|
||||||
|
renormalised for you. This is the R2 rule (client-rendered on-page) and the
|
||||||
|
I1 rule (unauditable off-page), finally computed instead of done by hand.
|
||||||
|
**N/A is not a zero** and the engine will not let it behave like one.
|
||||||
|
- `status: "error"` → malformed findings. Fix them; never fall back to
|
||||||
|
eyeballing a number.
|
||||||
|
- Run it twice on the same file before publishing. If the output moved, your
|
||||||
|
findings moved, and that is the thing to explain.
|
||||||
|
|
||||||
**Technical axis note:** CWV scored on CrUX field data (75th percentile,
|
**Technical axis note:** CWV scored on CrUX field data (75th percentile,
|
||||||
real users, from STEP 4) when available; otherwise lab PageSpeed
|
real users, from STEP 4) when available; otherwise lab PageSpeed
|
||||||
Lighthouse run.
|
Lighthouse run.
|
||||||
|
|||||||
@@ -193,6 +193,28 @@ fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500]
|
|||||||
• Mock is pages.json ({url: html}), not a single page.html: one fixture
|
• Mock is pages.json ({url: html}), not a single page.html: one fixture
|
||||||
cannot express a graph — every node would carry identical links.
|
cannot express a graph — every node would carry identical links.
|
||||||
|
|
||||||
|
fetch.sh score --findings <path.json | ->
|
||||||
|
→ {"status":"ok","axes":{"technical":{"score_20":17.8,"weight":0.2,
|
||||||
|
"weight_renormalised":0.2857,"findings":2}},
|
||||||
|
"na":["off-page","on-page"],"weights_renormalised":true,"global_20":17.6}
|
||||||
|
→ {"status":"error","reason":"unknown severity: 'bogus'"|"bad_findings_json"}
|
||||||
|
|
||||||
|
I7. /harden has a real scale (SKILL.md:435: -15/-8/-3/-1, clamp [0,100]);
|
||||||
|
/seo had none, so every axis was FELT and two runs over identical code could
|
||||||
|
disagree — while /client-handover gates on 17/20. Same scale here, /5 into
|
||||||
|
/20, one vocabulary across the family.
|
||||||
|
• The split: WHICH findings exist and how severe each is stays the LLM's
|
||||||
|
judgement. The addition is not. Same findings in, same score out.
|
||||||
|
• affected/sampled shift severity ONE step: >=50% of the sample escalates,
|
||||||
|
a single page de-escalates. A defect on 1 of 12 pages is not the defect
|
||||||
|
on 12 of 12.
|
||||||
|
• status:"na" → axis EXCLUDED, remaining weights renormalised. This is
|
||||||
|
R2's rule (client-rendered on-page) and I1's (unauditable off-page),
|
||||||
|
computed rather than done by hand. N/A is not a zero, and the engine
|
||||||
|
will not let it act like one.
|
||||||
|
• Malformed input is an error, never a silently wrong number — unlike the
|
||||||
|
fetch verbs, a degrade here would mean bad input, not a network fact.
|
||||||
|
|
||||||
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
|
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
|
||||||
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
|
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
|
||||||
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
|
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
|
||||||
|
|||||||
@@ -32,6 +32,8 @@ case "$cmd" in
|
|||||||
# No auth, no Google: stdlib-only, runs even without the venv.
|
# No auth, no Google: stdlib-only, runs even without the venv.
|
||||||
sitemap)
|
sitemap)
|
||||||
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
|
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
|
||||||
|
score)
|
||||||
|
exec "$PY" "$HERE/score.py" --store "$STORE" "$@" ;;
|
||||||
drift)
|
drift)
|
||||||
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
|
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
|
||||||
rendercheck)
|
rendercheck)
|
||||||
@@ -50,6 +52,6 @@ case "$cmd" in
|
|||||||
fi
|
fi
|
||||||
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
||||||
exit 2 ;;
|
exit 2 ;;
|
||||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|forget} [flags]"}'
|
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|forget} [flags]"}'
|
||||||
exit 2 ;;
|
exit 2 ;;
|
||||||
esac
|
esac
|
||||||
|
|||||||
@@ -0,0 +1,113 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Deterministic /20 scoring from a findings list. Stdlib only.
|
||||||
|
|
||||||
|
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
|
||||||
|
Basse -1, clamp [0,100]). /seo has none: every axis is felt, not computed, so
|
||||||
|
two runs over identical code can produce different scores. That is a
|
||||||
|
credibility problem on its own, and /client-handover gates on 17/20 — a
|
||||||
|
wobbling number makes the gate arbitrary. H2 sharpens it further: now that
|
||||||
|
drift reports what actually changed, a score moving on its own is visibly
|
||||||
|
noise.
|
||||||
|
|
||||||
|
The split is the point. The LLM keeps the irreducible judgement — WHICH
|
||||||
|
findings exist and how severe each is. The arithmetic stops being judgement:
|
||||||
|
same findings in, same score out. Same principle as grouping cannibalisation
|
||||||
|
rows in the engine rather than asking a model to add up 1000 of them.
|
||||||
|
|
||||||
|
Scale is /harden's, /5 into /20, so the whole skill family speaks one
|
||||||
|
vocabulary.
|
||||||
|
"""
|
||||||
|
import argparse, json, sys
|
||||||
|
|
||||||
|
PENALTY = {"critique": 15, "haute": 8, "moyenne": 3, "basse": 1}
|
||||||
|
|
||||||
|
# STEP 9 weights. FULL = 7 axes, LOCAL = 4 (off-page/social/competitive are
|
||||||
|
# not audited at that depth).
|
||||||
|
WEIGHTS = {
|
||||||
|
("FULL", "local"): {"technical": .20, "on-page": .20, "seo-local": .25,
|
||||||
|
"off-page": .10, "social": .10, "competitive": .05,
|
||||||
|
"legal": .10},
|
||||||
|
("FULL", "national"): {"technical": .30, "on-page": .30, "seo-local": .05,
|
||||||
|
"off-page": .15, "social": .05, "competitive": .10,
|
||||||
|
"legal": .05},
|
||||||
|
("LOCAL", "local"): {"technical": .25, "on-page": .35, "seo-local": .20,
|
||||||
|
"legal": .20},
|
||||||
|
("LOCAL", "national"):{"technical": .35, "on-page": .45, "seo-local": .05,
|
||||||
|
"legal": .15},
|
||||||
|
}
|
||||||
|
|
||||||
|
def _axis_score(findings):
|
||||||
|
"""100 - Σ penalties, clamped, then /5 → /20. Prevalence shifts severity
|
||||||
|
ONE step, never invents one: a finding on 1 of 12 sampled pages is not the
|
||||||
|
same defect as one on 12 of 12, and pretending otherwise is what made the
|
||||||
|
old scores unreproducible."""
|
||||||
|
total = 0
|
||||||
|
for f in findings:
|
||||||
|
sev = str(f.get("severity", "")).lower()
|
||||||
|
if sev not in PENALTY:
|
||||||
|
raise ValueError("unknown severity: %r" % f.get("severity"))
|
||||||
|
order = ["basse", "moyenne", "haute", "critique"]
|
||||||
|
i = order.index(sev)
|
||||||
|
aff, samp = f.get("affected"), f.get("sampled")
|
||||||
|
if isinstance(aff, int) and isinstance(samp, int) and samp > 0:
|
||||||
|
ratio = aff / samp
|
||||||
|
if ratio >= 0.5:
|
||||||
|
i = min(i + 1, len(order) - 1) # widespread → escalate
|
||||||
|
elif aff <= 1:
|
||||||
|
i = max(i - 1, 0) # isolated → de-escalate
|
||||||
|
total += PENALTY[order[i]]
|
||||||
|
return round(max(0, 100 - total) / 5.0, 1)
|
||||||
|
|
||||||
|
def score(payload):
|
||||||
|
depth = str(payload.get("depth", "FULL")).upper()
|
||||||
|
profile = str(payload.get("profile", "local")).lower()
|
||||||
|
key = (depth, profile)
|
||||||
|
if key not in WEIGHTS:
|
||||||
|
return {"status": "error", "reason": "unknown depth/profile: %s/%s"
|
||||||
|
% (depth, profile)}
|
||||||
|
weights, axes_in = WEIGHTS[key], payload.get("axes", {})
|
||||||
|
scored, na = {}, []
|
||||||
|
for axis, w in weights.items():
|
||||||
|
a = axes_in.get(axis)
|
||||||
|
if a is None or str(a.get("status", "")).lower() == "na":
|
||||||
|
na.append(axis) # N/A is not a zero
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
s = _axis_score(a.get("findings", []))
|
||||||
|
except ValueError as e:
|
||||||
|
return {"status": "error", "reason": str(e)}
|
||||||
|
scored[axis] = {"score_20": s, "weight": w,
|
||||||
|
"findings": len(a.get("findings", []))}
|
||||||
|
if not scored:
|
||||||
|
return {"status": "degraded", "reason": "no_axis_scored"}
|
||||||
|
# Renormalise over what was actually measured. R2 mandates this for a
|
||||||
|
# client-rendered on-page axis and left it to the model to do by hand.
|
||||||
|
live = sum(v["weight"] for v in scored.values())
|
||||||
|
for v in scored.values():
|
||||||
|
v["weight_renormalised"] = round(v["weight"] / live, 4)
|
||||||
|
glob = sum(v["score_20"] * v["weight"] / live for v in scored.values())
|
||||||
|
return {"status": "ok", "source": "score", "depth": depth,
|
||||||
|
"profile": profile, "axes": scored, "na": sorted(na),
|
||||||
|
"weights_renormalised": round(live, 4) != 1.0,
|
||||||
|
"global_20": round(glob, 1)}
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
try:
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
p.add_argument("--findings", default="-", help="JSON path, or - for stdin")
|
||||||
|
p.add_argument("--store", default=None) # accepted+ignored
|
||||||
|
args = p.parse_args()
|
||||||
|
raw = sys.stdin.read() if args.findings == "-" else \
|
||||||
|
open(args.findings, encoding="utf-8").read()
|
||||||
|
print(json.dumps(score(json.loads(raw)), indent=2))
|
||||||
|
except SystemExit as e:
|
||||||
|
if e.code not in (0, None):
|
||||||
|
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||||
|
raise
|
||||||
|
except Exception:
|
||||||
|
# Unlike the fetch verbs this is pure arithmetic: a degrade here means
|
||||||
|
# malformed input, never a network fact.
|
||||||
|
print(json.dumps({"status": "error", "reason": "bad_findings_json"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
@@ -178,6 +178,35 @@ has "cap is reported" "$CAP" '"capped": true'
|
|||||||
has "capped withholds orphans" "$CAP" '"orphans_withheld": true'
|
has "capped withholds orphans" "$CAP" '"orphans_withheld": true'
|
||||||
hasnt "capped emits no orphans" "$CAP" '"orphans":'
|
hasnt "capped emits no orphans" "$CAP" '"orphans":'
|
||||||
|
|
||||||
|
echo "── score (I7) ──"
|
||||||
|
sc() { printf '%s' "$1" | python3 "$SD/score.py" --findings -; }
|
||||||
|
# technical: haute(-8) + moyenne(-3) = 100-11 = 89 → 17.8
|
||||||
|
B='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"haute"},{"severity":"moyenne"}]},"seo-local":{"findings":[]},"off-page":{"findings":[]},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]},"on-page":{"findings":[]}}}'
|
||||||
|
R="$(sc "$B")"
|
||||||
|
has "harden scale, /5 into /20" "$R" '"score_20": 17.8'
|
||||||
|
has "no findings = 20" "$R" '"score_20": 20.0'
|
||||||
|
has "nothing renormalised" "$R" '"weights_renormalised": false'
|
||||||
|
# THE point of I7: same findings in, same score out
|
||||||
|
A1="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
|
||||||
|
A2="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
|
||||||
|
[ "$A1" = "$A2" ] && ok "score is reproducible" || no "score is reproducible" "$A1 vs $A2"
|
||||||
|
# N/A is not a zero, and R2 mandated renormalising by hand — now computed
|
||||||
|
NA='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[]},"on-page":{"status":"na"},"seo-local":{"findings":[]},"off-page":{"status":"na"},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]}}}'
|
||||||
|
RN="$(sc "$NA")"
|
||||||
|
has "na axes listed" "$RN" '"on-page"'
|
||||||
|
has "renormalisation flagged" "$RN" '"weights_renormalised": true'
|
||||||
|
# all axes 20 → global must stay 20: N/A must not drag the mean down
|
||||||
|
has "na is not a zero" "$RN" '"global_20": 20.0'
|
||||||
|
# prevalence shifts severity ONE step, both ways
|
||||||
|
WIDE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":10,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
|
||||||
|
ONE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":1,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
|
||||||
|
has "widespread escalates (-8)" "$(sc "$WIDE")" '"score_20": 18.4'
|
||||||
|
has "isolated de-escalates (-1)" "$(sc "$ONE")" '"score_20": 19.8'
|
||||||
|
# malformed input is an error, never a silently wrong number
|
||||||
|
has "unknown severity rejected" "$(sc '{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"bogus"}]}}}')" '"status": "error"'
|
||||||
|
has "unknown profile rejected" "$(sc '{"depth":"FULL","profile":"martian","axes":{}}')" '"status": "error"'
|
||||||
|
has "garbage json is an error" "$(sc 'not json')" '"status": "error"'
|
||||||
|
|
||||||
echo "── drift (H2) ──"
|
echo "── drift (H2) ──"
|
||||||
DH="$(mktemp -d)"
|
DH="$(mktemp -d)"
|
||||||
D1="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v1" python3 "$SD/drift.py" \
|
D1="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v1" python3 "$SD/drift.py" \
|
||||||
|
|||||||
Reference in New Issue
Block a user