feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3, Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over identical code could disagree. That is a credibility problem on its own, and /client-handover gates on 17/20 — a wobbling number makes the gate arbitrary. H2 sharpened it: now that drift reports what actually changed, a score moving on its own is visibly noise. The split is the whole point. WHICH findings exist and how severe each is stays the LLM's judgement — irreducible, and I am not pretending otherwise. The arithmetic stops being judgement: same findings in, same score out. Same principle as grouping cannibalisation rows in the engine rather than handing a model 1000 rows to add up. Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary instead of two. Two things it makes real that were prose: - **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable off-page) both mandate excluding an axis and renormalising the rest. Both left that arithmetic to the model. Now the engine does it and refuses to let N/A behave like a zero — verified: all-20 axes with two N/A still yields global 20.0, not a dragged-down mean. - **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a single page de-escalates). A defect on 1 of 12 pages is not the defect on 12 of 12, and flattening the two is part of what made the old numbers move. Malformed input is an error, never a silently wrong number — unlike the fetch verbs, a degrade here would mean bad input, not a network fact. Unknown severity and unknown profile both rejected, tested. Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8; critique+haute = 77 → 15.4), identical global across repeated runs, weights renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail; full suite green; shellcheck + py_compile clean.
This commit is contained in:
@@ -178,6 +178,35 @@ has "cap is reported" "$CAP" '"capped": true'
|
||||
has "capped withholds orphans" "$CAP" '"orphans_withheld": true'
|
||||
hasnt "capped emits no orphans" "$CAP" '"orphans":'
|
||||
|
||||
echo "── score (I7) ──"
|
||||
sc() { printf '%s' "$1" | python3 "$SD/score.py" --findings -; }
|
||||
# technical: haute(-8) + moyenne(-3) = 100-11 = 89 → 17.8
|
||||
B='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"haute"},{"severity":"moyenne"}]},"seo-local":{"findings":[]},"off-page":{"findings":[]},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]},"on-page":{"findings":[]}}}'
|
||||
R="$(sc "$B")"
|
||||
has "harden scale, /5 into /20" "$R" '"score_20": 17.8'
|
||||
has "no findings = 20" "$R" '"score_20": 20.0'
|
||||
has "nothing renormalised" "$R" '"weights_renormalised": false'
|
||||
# THE point of I7: same findings in, same score out
|
||||
A1="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
|
||||
A2="$(sc "$B" | python3 -c 'import sys,json;print(json.load(sys.stdin)["global_20"])')"
|
||||
[ "$A1" = "$A2" ] && ok "score is reproducible" || no "score is reproducible" "$A1 vs $A2"
|
||||
# N/A is not a zero, and R2 mandated renormalising by hand — now computed
|
||||
NA='{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[]},"on-page":{"status":"na"},"seo-local":{"findings":[]},"off-page":{"status":"na"},"social":{"findings":[]},"competitive":{"findings":[]},"legal":{"findings":[]}}}'
|
||||
RN="$(sc "$NA")"
|
||||
has "na axes listed" "$RN" '"on-page"'
|
||||
has "renormalisation flagged" "$RN" '"weights_renormalised": true'
|
||||
# all axes 20 → global must stay 20: N/A must not drag the mean down
|
||||
has "na is not a zero" "$RN" '"global_20": 20.0'
|
||||
# prevalence shifts severity ONE step, both ways
|
||||
WIDE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":10,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
|
||||
ONE='{"depth":"LOCAL","profile":"local","axes":{"technical":{"findings":[{"severity":"moyenne","affected":1,"sampled":12}]},"on-page":{"findings":[]},"seo-local":{"findings":[]},"legal":{"findings":[]}}}'
|
||||
has "widespread escalates (-8)" "$(sc "$WIDE")" '"score_20": 18.4'
|
||||
has "isolated de-escalates (-1)" "$(sc "$ONE")" '"score_20": 19.8'
|
||||
# malformed input is an error, never a silently wrong number
|
||||
has "unknown severity rejected" "$(sc '{"depth":"FULL","profile":"local","axes":{"technical":{"findings":[{"severity":"bogus"}]}}}')" '"status": "error"'
|
||||
has "unknown profile rejected" "$(sc '{"depth":"FULL","profile":"martian","axes":{}}')" '"status": "error"'
|
||||
has "garbage json is an error" "$(sc 'not json')" '"status": "error"'
|
||||
|
||||
echo "── drift (H2) ──"
|
||||
DH="$(mktemp -d)"
|
||||
D1="$(HOME="$DH" SEO_DATA_MOCK_DIR="$SD/fixtures-drift-v1" python3 "$SD/drift.py" \
|
||||
|
||||
Reference in New Issue
Block a user