feat(seo-data): content_quality verb — deterministic filler/AI-slop signal

Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
content_quality.py, rewritten to the lib/seo-data contract per BDR-070. The
Content Shape axis was 100% LLM judgement; this gives it a measured input.

fetch.sh content_quality (stdin or --file) → {filler_score, ai_pattern_score,
information_density, overall_quality, flags[], matches{}}. 100% deterministic:
QRG §4.6 filler list (26 phrases) + AI-pattern list (46) kept intact, regex
matching, no LLM. Stdlib only (argparse/json/re/sys/collections/typing).

Advisory, NOT a verdict — the point of the wiring. It never claims a page "is
AI-written" (LRN-131/133); flags are candidates for human review. geo-analyzer
STEP 8 Check 10 makes it a deterministic input that INFORMS checks 1-9, never
replaces them, never scored on its own. A low number is not an automatic
finding.

Detection proven both directions (a detector that always- or never-flags is
useless): filler+slop text → flags [filler, low-density], overall 34-49; clean
dense factual text (dates/EUR/percentages) → no flags, overall 90. Empty input →
degraded/empty_input, never zeros-as-a-result.

Verified: GATE 1 verifier CONFORME 10/10 (both directions exercised live, lists
diffed intact vs source, advisory language confirmed); GATE 2 self-scan clean
(only sink is read-only open() for --file); seo-data 190 → 210 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 19:06:58 +02:00
parent fb0b587240
commit b271e83fb6
5 changed files with 371 additions and 1 deletions
+15
View File
@@ -551,6 +551,13 @@ sample of a 300-page site says nothing about the other 294.
pronouns?
8. **Lists/tables vs prose** — structured where possible?
9. **30/70 rule** (if city/service variants exist) — ≥70% unique?
10. **Filler/AI-slop signal (deterministic)** — feed each sampled page's
body text to `fetch.sh content_quality`. It is a DETERMINISTIC input
that INFORMS checks 1-9 (word-list/density heuristics, no LLM call);
it never replaces your read of them. A low `overall_quality` or a
`filler`/`ai-patterns` flag is a candidate for human review, not an
automatic finding — do not let the number become the verdict, and do
not claim a page "is AI-written" from it.
### Sampling command
@@ -561,6 +568,11 @@ for f in index.html $(find . "${FEXCL[@]}" -maxdepth 3 \( -name "*.astro" -o -na
echo "=== $f ==="
grep -oE '<(h1|h2|h3)[^>]*>[^<]+</(h1|h2|h3)>|^#{1,3} .+' "$f" 2>/dev/null | head -20
done
# Filler/AI-slop signal (Check 10) — strip markup to plain body text, then
# score it. Advisory only: pair the number with your own read of Checks 1-9.
sed -e 's/<[^>]*>//g' index.html | \
bash ~/.claude/lib/seo-data/fetch.sh content_quality
```
### Findings
@@ -576,6 +588,9 @@ CITED STATISTICS : <avg per page>
FRESHNESS VISIBLE : <n/N pages>
PRONOUN-HEAVY : <n/N pages flagged>
30/70 RULE : pass | fail | N/A
FILLER/AI-SLOP SIGNAL : <avg overall_quality>/100, flags: <n/N pages flagged>
(deterministic, advisory — informs checks 1-9, never
a verdict, never scored on its own)
PRIORITY ACTIONS : <top 5>
```