forked from bchanot/claude
feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3, Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over identical code could disagree. That is a credibility problem on its own, and /client-handover gates on 17/20 — a wobbling number makes the gate arbitrary. H2 sharpened it: now that drift reports what actually changed, a score moving on its own is visibly noise. The split is the whole point. WHICH findings exist and how severe each is stays the LLM's judgement — irreducible, and I am not pretending otherwise. The arithmetic stops being judgement: same findings in, same score out. Same principle as grouping cannibalisation rows in the engine rather than handing a model 1000 rows to add up. Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary instead of two. Two things it makes real that were prose: - **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable off-page) both mandate excluding an axis and renormalising the rest. Both left that arithmetic to the model. Now the engine does it and refuses to let N/A behave like a zero — verified: all-20 axes with two N/A still yields global 20.0, not a dragged-down mean. - **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a single page de-escalates). A defect on 1 of 12 pages is not the defect on 12 of 12, and flattening the two is part of what made the old numbers move. Malformed input is an error, never a silently wrong number — unlike the fetch verbs, a degrade here would mean bad input, not a network fact. Unknown severity and unknown profile both rejected, tested. Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8; critique+haute = 77 → 15.4), identical global across repeated runs, weights renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail; full suite green; shellcheck + py_compile clean.
This commit is contained in:
@@ -193,6 +193,28 @@ fetch.sh linkgraph --url https://ex.com/sitemap.xml [--max 500]
|
||||
• Mock is pages.json ({url: html}), not a single page.html: one fixture
|
||||
cannot express a graph — every node would carry identical links.
|
||||
|
||||
fetch.sh score --findings <path.json | ->
|
||||
→ {"status":"ok","axes":{"technical":{"score_20":17.8,"weight":0.2,
|
||||
"weight_renormalised":0.2857,"findings":2}},
|
||||
"na":["off-page","on-page"],"weights_renormalised":true,"global_20":17.6}
|
||||
→ {"status":"error","reason":"unknown severity: 'bogus'"|"bad_findings_json"}
|
||||
|
||||
I7. /harden has a real scale (SKILL.md:435: -15/-8/-3/-1, clamp [0,100]);
|
||||
/seo had none, so every axis was FELT and two runs over identical code could
|
||||
disagree — while /client-handover gates on 17/20. Same scale here, /5 into
|
||||
/20, one vocabulary across the family.
|
||||
• The split: WHICH findings exist and how severe each is stays the LLM's
|
||||
judgement. The addition is not. Same findings in, same score out.
|
||||
• affected/sampled shift severity ONE step: >=50% of the sample escalates,
|
||||
a single page de-escalates. A defect on 1 of 12 pages is not the defect
|
||||
on 12 of 12.
|
||||
• status:"na" → axis EXCLUDED, remaining weights renormalised. This is
|
||||
R2's rule (client-rendered on-page) and I1's (unauditable off-page),
|
||||
computed rather than done by hand. N/A is not a zero, and the engine
|
||||
will not let it act like one.
|
||||
• Malformed input is an error, never a silently wrong number — unlike the
|
||||
fetch verbs, a degrade here would mean bad input, not a network fact.
|
||||
|
||||
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
|
||||
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
|
||||
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
|
||||
|
||||
Reference in New Issue
Block a user