Commit Graph
726 Commits
Author SHA1 Message Date
Bastien Chanot d2a10de08b feat(agents): W4 handover two-mode — synthesize opus / render sonnet, run-scoped draft handoff (BDR-077)
handover-doc-writer: MODE synthesize (model="opus" call-site — STEP
9/10/12 → .audit/handover-draft-<RUNID>.md + DRAFT COMPLETE sentinel) /
MODE render (sonnet pin — STEP 13-16 from the draft, fail-closed on
absent/mismatched RUNID). Mode-based, not a file split: the §9
name+dispatch census locks survive untouched. client-handover-writer 9.6
dispatches twice with the FULL PACKAGE both times (LRN-126) + RUNID mint
+ post-run draft cleanup. In-wave smokes PASSED disk-verified: draft
written+sentinel+gitignored; render BLOCKED on wrong RUNID (no phantom
synthesis); render consumed draft + honored skip-write, report grammar
intact. Census §17 (111 pass), make test exit 0.
2026-07-19 22:59:34 +02:00
Bastien Chanot d3d5e3802c Merge feature/model-tiering-w3-tier-moves into develop 2026-07-19 22:46:19 +02:00
Bastien Chanot 5e8bb0c22e feat(routing): W3 tier moves — validator-analyzer opus→sonnet, commit-changer propose=opus override (BDR-077)
validator-analyzer tiered down (deterministic validator-runner + fixed
deduction tables — no deep judgment; supersedes its BDR-076 opus pin).
commit-changer: MODE propose dispatched model="opus" (narrative
reconstruction + capitalize routing = judgment), MODE apply on the sonnet
pin. Typed-pin precedence smoke PASSED: sonnet-pinned verifier dispatched
model="haiku" ran on claude-haiku-4-5 — call-site wins, documented +
now behaviorally proven. Census §11 flip + §16 (106 pass), make test
exit 0.
2026-07-19 22:45:00 +02:00
Bastien Chanot e4d2629c88 Merge feature/model-tiering-w2-conversions into develop 2026-07-19 22:42:39 +02:00
Bastien Chanot 18075a38db feat(agents): W2/S2 doc pipeline two-mode + last inline conversions (BDR-077)
doc-syncer: ONE agent, TWO dispatch modes around the dispatcher's gate —
MODE: audit (model="opus" call-site override, READ-ONLY, drafts + PATCH
PLAN) / MODE: patch (sonnet pin, applies the APPROVED plan, shape oracle
w/ revert-on-fail, emits CHANGE SUMMARY + PATCHED_FILES). Deviation from
plan's 2-file split, per the challenge's own commit-changer mode
precedent: zero text duplication, zero lock moves. Fixes a LATENT DEFECT:
/doc dispatched an agent whose STEP 8 gate could never fire (dispatched
agents cannot ask) — the gate now lives in the dispatcher (DISPATCHER
PROTOCOL section). doc-commit.md consumes the patcher's CHANGE SUMMARY
(the in-thread context now crosses the dispatch boundary, LRN-126).
Consumers rewired: /doc (audit→gate→patch→commit), onboard (audit
report-only, opus), doc-commit steps in bugfix/hotfix/feat/ship-feature/
init-project(5b+10c); scaffolder loses PHASE 6 (README = init 5b's job);
scaffolder + onboarder now DISPATCHED in init-project/onboard (pins live,
was inline on session model). In-wave planted-drift smoke PASSED
end-to-end, disk-verified (audit caught npm-run-dev drift → [MINOR] plan
→ patch applied → summary crossed). Census §14-15 (103 pass), make test
exit 0. Typed plugin-probe dispatch resolution verified post-restart.
2026-07-19 22:28:09 +02:00
Bastien Chanot 74528a6910 feat(agents): W2/S1 plugin split — probe (sonnet) + advisor reasoner (opus) + plugin-gate include (BDR-077)
plugin-advisor keeps its name, becomes the opus REASONER: PHASE 1 bash
extracted to new plugin-probe (sonnet, facts-only PROBE REPORT), PHASE 4
apply + checkpoint hoisted to new lib/plugin-gate.md (main-loop include,
doc-commit.md x6 pattern). Fail-closed: advisor ERRORs on missing report.
4 consumers rewired (plugin-check, onboard, init-project, ship-feature).
In-wave planted-input smoke PASSED: probe report complete w/ fallbacks;
advisor consumed every planted field (monorepo per-package note, fast-libs
ctx7 reco) with zero re-detection; ERROR verdict on absent report.
Census §13 (81 pass). Note: new subagent_type registers next session —
resolution re-check before wave merge.
2026-07-19 20:57:10 +02:00
Bastien Chanot 896d3faaf9 Merge feature/model-tiering-w1-no-inherit into develop 2026-07-19 20:00:37 +02:00
Bastien Chanot 3f7c754239 feat(routing): W1 no-inherit — fable skill-runners, opus review dispatches, dispatch-tier doctrine (BDR-077)
No dispatched agent inherits the session model anymore:
- client-handover-writer's 7 general-purpose skill-runner dispatch sites
  carry model: "fable" (+ normative rule; spike-verified alias — resolves
  claude-fable-5, enum-validated, loud failure, never silent fallback)
- ship-feature STEP 6 + init-project STEP 10 code-review dispatches carry
  model: "opus" (was: inherit — the leak the maps exposed)
- model-gate.md §4: dispatch-tier doctrine (typed = frontmatter pin,
  built-ins = explicit model= at every call site)
- census §12 (66 pass), make test green
2026-07-19 19:51:06 +02:00
Bastien Chanot 1c2d30dbf0 chore(tasks): model-tiering v2 — analysis + challenged plan v3 + TODO reconcile
Plan challenged by 3 blind lenses + 1 confirmation pass (1 BLOCKER closed
by fable-dispatch spike, 6 MAJORs + 8 MINORs fixed by named changes, 0
deferred). TODO: seo-geo-integrity 'UNMERGED' note was stale (92301fe
already in develop) — corrected.
2026-07-19 19:51:06 +02:00
Bastien Chanot 17fbe51aa1 Merge feature/opus-pin-audit-agents into develop 2026-07-19 19:46:10 +02:00
Bastien Chanot 3eaf31ca09 chore(memory): BDR-076 + journal + TODO — opus-pin dispatched judgment agents 2026-07-19 17:38:58 +02:00
Bastien Chanot 354ff2644f feat(agents): pin dispatched judgment agents to opus — Fable = inline reflection only (BDR-076)
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
2026-07-19 17:38:55 +02:00
Bastien Chanot 9bc6ab7e07 Merge feature/hotfix-challenge-guard into develop 2026-07-18 23:24:04 +02:00
Bastien Chanot 727a41ad71 chore(memory): BDR-075 amendment (hotfix included) + journal — Option B + behavioral smoke 2026-07-18 23:08:53 +02:00
Bastien Chanot 311ea14789 feat(skills): wire hotfix into plan-challenge under a logic-only guard
STEP 1.8 (Option B): skip purely cosmetic fixes (CSS/copy/typo), run the 3-lens
challenge only when the fix touches control flow/behaviour (off-by-one, wrong
operator, behaviour-changing config, execution-altering import); a BLOCKER means
it was never a hotfix -> escalate to /bugfix. 12th orchestrator wired.

- skills/hotfix/SKILL.md          — STEP 1.8 guarded challenge
- lib/tests/plan-challenger.test.sh — lock hotfix into the census (43 assertions)
2026-07-18 23:08:52 +02:00
Bastien Chanot a68f26ca9c Merge feature/plan-challenge-phase into develop 2026-07-17 22:57:40 +02:00
Bastien Chanot d5d1584c1c Merge feature/drop-config-protection into develop 2026-07-17 22:57:10 +02:00
Bastien Chanot 2aa95636ee chore(memory): BDR-074 BDR-075 EVAL-026 LRN-136 — config-protection removal + plan-challenge phase 2026-07-17 22:54:12 +02:00
Bastien Chanot 6bfc0543e5 feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.

- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
  severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
  onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)

Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
2026-07-17 22:51:50 +02:00
Bastien Chanot 0e1b89c71a chore(hooks): remove config-protection edit-block guardrail
Full removal per user request: the PreToolUse hook that blocked model
Edit/Write on quality-gate files (settings.json, gitflow.sh, .githooks,
doctor.sh, hooks, lib/tests, lint configs) plus its one-shot sentinel.

- delete hooks/config-protection.sh
- delete lib/tests/config-protection.test.sh
- deregister the hook from settings.json (rtk-rewrite PreToolUse kept)
- drop the README mention

Residual protection unchanged: gitflow pre-commit guard + Gitea branch
protection still block direct code commits to main/develop.
2026-07-17 21:56:32 +02:00
Bastien Chanot 23c8c290d7 Merge feature/dns-rebinding-guard into develop 2026-07-17 20:17:45 +02:00
Bastien Chanot a391be4906 chore(memory): LRN-134 LRN-135 — capitalize 2026-07-17 20:16:50 +02:00
Bastien Chanot b00e8ef442 feat(seo-data): safe_fetch — resolve-then-pin, close DNS-rebinding + SSRF
By-principle hardening. H1's url-guard validates the NAME; urlopen then resolved
AND connected — two DNS lookups with a window a hostile authority uses to answer
PUBLIC to validation and PRIVATE (169.254.169.254 metadata, 127.0.0.1, the LAN)
to the connect. A name-level guard cannot see that rebind.

safe_fetch collapses the two lookups into one: resolve ONCE, validate every IP
(ipaddress, dual-stack v4+v6), refuse if ANY is non-public (the multi-A vector),
connect to the exact validated IP with Host+SNI+cert for the real host — no
second resolution to poison. Redirects re-validate each hop (urlopen followed
them blind). One seam: sitemap._fetch, which linkgraph/render_check/drift all
call, so every network verb inherits it.

The load-bearing property (confirmed by the security review): classification is
on the OS-resolved address (sockaddr[0]), never the URL text — so octal/hex/
decimal literals, IPv4-mapped IPv6, NAT64, 6to4 are all defeated structurally,
not by enumeration.

Better than the source idea (claude-seo url_safety.py, MIT): dual-stack (theirs
IPv4-only), no global monkeypatch so thread-safe by construction (theirs locks a
patched getaddrinfo), stdlib-only (no requests). Proven end-to-end before
writing: pinned connect keeps SNI+cert for the real host.

NOT covered, stated not silent: shell `curl` in the agent specs (separate
process, unpinnable here). Smaller surface; `curl --resolve` is a separate change.

REVIEW-SURFACED (fresh security-auditor, adversarial, VERDICT PASS) — two real
holes it found while attacking the diff, both fixed here:
- billion-laughs REOPENED in C1b: _refuse_dtd scanned only raw[:4096], so a
  >4KB leading comment pushed <!DOCTYPE past the window while ET parsed AND
  EXPANDED the entities. Proven (&lol2; → "lollollollollol"), now a full-doc
  case-insensitive scan. This is a genuine fix to already-merged C1b, not this
  feature — fixed here rather than filed, per root-cause discipline.
- 192.88.99.0/24 (6to4-relay anycast) passed is_global as public — added to an
  extra special-use deny list.

Verified: rebind-to-metadata refused BEFORE any connect (injected resolver),
multi-A public+private refused, classifier fuzzed dual-stack incl. CGNAT/6to4,
non-http scheme refused, both review fixes proven with no false positive; real
fetch still works (zenquality 86 loc, lavageangels 24) through the pinned path;
all 4 verbs work end-to-end via fetch.sh; seo-data 210 → 221 pass, 0 fail; full
suite green; shellcheck + py_compile clean.
2026-07-17 19:58:59 +02:00
Bastien Chanot 4ccfb606a8 Merge feature/seo-data-cherry-picks into develop 2026-07-17 19:10:10 +02:00
Bastien Chanot 0564afcb3c chore(memory): journal — content_quality shipped, both easy picks done 2026-07-17 19:07:10 +02:00
Bastien Chanot b271e83fb6 feat(seo-data): content_quality verb — deterministic filler/AI-slop signal
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
content_quality.py, rewritten to the lib/seo-data contract per BDR-070. The
Content Shape axis was 100% LLM judgement; this gives it a measured input.

fetch.sh content_quality (stdin or --file) → {filler_score, ai_pattern_score,
information_density, overall_quality, flags[], matches{}}. 100% deterministic:
QRG §4.6 filler list (26 phrases) + AI-pattern list (46) kept intact, regex
matching, no LLM. Stdlib only (argparse/json/re/sys/collections/typing).

Advisory, NOT a verdict — the point of the wiring. It never claims a page "is
AI-written" (LRN-131/133); flags are candidates for human review. geo-analyzer
STEP 8 Check 10 makes it a deterministic input that INFORMS checks 1-9, never
replaces them, never scored on its own. A low number is not an automatic
finding.

Detection proven both directions (a detector that always- or never-flags is
useless): filler+slop text → flags [filler, low-density], overall 34-49; clean
dense factual text (dates/EUR/percentages) → no flags, overall 90. Empty input →
degraded/empty_input, never zeros-as-a-result.

Verified: GATE 1 verifier CONFORME 10/10 (both directions exercised live, lists
diffed intact vs source, advisory language confirmed); GATE 2 self-scan clean
(only sink is read-only open() for --file); seo-data 190 → 210 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 19:06:58 +02:00
Bastien Chanot fb0b587240 chore(memory): journal — schema_gen shipped, gap-revisit note 2026-07-17 14:31:26 +02:00
Bastien Chanot cfdd89e73b feat(seo-data): schema_gen verb — generate JSON-LD, not just audit it
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.

fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.

Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).

geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.

Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 14:30:31 +02:00
Bastien Chanot 92301fe1c8 Merge bugfix/seo-geo-integrity into develop 2026-07-17 13:48:13 +02:00
Bastien Chanot f96206ff21 chore(memory): BDR-070..073 LRN-131..133 BLK-017 EVAL-025 — capitalize 2026-07-17 13:46:17 +02:00
Bastien Chanot 4818c6116f feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.

The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.

Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.

Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
  off-page) both mandate excluding an axis and renormalising the rest. Both
  left that arithmetic to the model. Now the engine does it and refuses to let
  N/A behave like a zero — verified: all-20 axes with two N/A still yields
  global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
  single page de-escalates). A defect on 1 of 12 pages is not the defect on
  12 of 12, and flattening the two is part of what made the old numbers move.

Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.

Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 13:29:14 +02:00
Bastien Chanot f69cfc5cb4 feat(seo-data): H2 — drift baseline; regressions vs changes, not prose
seo-analyzer.md:1365 keeps history as "date + score + key changes" — prose the
LLM writes about its own previous prose. Lossy, unreproducible, and
machine-uncomparable, so "the redesign silently dropped 40 canonicals" is
invisible unless someone happens to notice.

drift snapshots title/description/canonical/robots/h1_count/jsonld_types per
URL and diffs them. Stdlib only, no auth.

The classification IS the feature: LOSING a signal is a regression, CHANGING
one is a change that may well be intended. The engine says which kind; the
agent judges. A reworded title is not an alert; an evaporated canonical is.

Runs over the WHOLE sitemap, never a sample — caught while designing: a drift
computed over a sample that changes between runs compares nothing.

NOT rank tracking. That is the common misread of this same feature elsewhere;
positions come from GSC `queries`. This is on-page regression detection.

Also caught in my own draft before testing: _capture reused
sm._mock("page.html"), the exact single-fixture flaw I had already fixed in
linkgraph — one fixture cannot express a multi-page snapshot, every URL would
read identical. Now pages.json, same convention.

Proved on a planted failure rather than a happy path — two clean sites would
look identical to a detector that always returns []:
  v1 -> v2: canonical lost on /a, h1 + jsonld lost on /, title reworded,
  /gone removed, /neuve added
  → 3 regressions, 1 change, gone/new both detected, title correctly NOT a
    regression.

Store is ~/.claude/seo-data/drift/<host>.json, 0700, written via os.replace so
a crash never leaves a half-written baseline; a corrupt store degrades to
"first run" instead of killing the audit.

Verified: seo-data 144 -> 155 pass, 0 fail; full suite green.
2026-07-17 13:25:45 +02:00
Bastien Chanot d6b8edc8ea fix(seo): B1 KILLED — Common Crawl backlinks measured, not assumed
The plan said Common Crawl was the free backlink source and the 70/100 cap
was therefore mandatory. Measured before building, and both premises die.

HEAD against data.commoncrawl.org, live:
  cc-main-2026-feb-mar-apr-domain-edges.txt.gz    17.3 GB   gzipped
  cc-main-2026-feb-mar-apr-domain-ranks.txt.gz     2.3 GB
  cc-main-2026-feb-mar-apr-domain-vertices.txt.gz  879 MB

Finding one domain's inbound links means scanning the edges file end to end,
per audit. That is not slow, it is non-viable — and abusive toward a
nonprofit serving the data free.

Worse, the reference implementation everyone points at
(claude-seo scripts/commoncrawl_graph.py:169) does this:

    max_compressed_bytes = 500 * 1024 * 1024   # 500 MiB safety cap
    if total_downloaded > max_compressed_bytes: break

500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source
ID — so it reads an arbitrary slice of source domains and reports whatever
backlinks happened to be in it, as a backlink profile, capped at "70/100
health". Nothing in the output says 3%. That is a random sample wearing a
measurement's clothes: the exact failure class this branch exists to remove,
and I was one step from copying it.

B2 dies with B1: nothing left to cap.

CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand
mentions only, backlinks + authority declared unauditable in §14 — is the
FINAL state, not a placeholder waiting for data. Corrected my own I1 text,
which pointed at Common Crawl as the "nearest free source": that sends a
future reader into a 17 GB dead end. The §14 line now records what was
measured and why no number beats a fabricated one.

Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free
source") was wrong for the same reason. The only free viable backlink source
is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on
the client's Bing account. That raises W2's value; it does not unblock it.

Verified: full suite green, seo-data 144 pass / 0 fail.
2026-07-17 13:17:18 +02:00
Bastien Chanot 02c7a6fe6d Merge bugfix/seo-geo-integrity-phase1 into develop 2026-07-17 13:09:08 +02:00
Bastien Chanot 20d3082542 feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright
Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
2026-07-17 12:44:43 +02:00
Bastien Chanot fe41986be9 feat(seo-data): C3 — internal link graph; orphans + click depth, measured
seo-analyzer.md:613 asks "Every important page reachable within 3 clicks?"
and :616 asks "Orphan pages (no inbound internal links)?". Neither ever had a
command — same shape as the sameAs check before W3. This is that command.

My earlier reservation ("costs a lot of network") was wrong and the
measurement killed it: 24 pages in 2.7s, 86 in 3.8s. Cheap enough to always
run on FULL.

EXHAUSTIVE OR NOTHING is the design constraint, not a nicety. Orphans cannot
be sampled: proving a page has no inbound link means having read every other
page. So when the crawl is capped or any page fails, orphans are WITHHELD —
`orphans_withheld: true` and no list. A false orphan ("page X has no inbound
links" when it does) sends a client fixing what is not broken; that is the
worst finding this tool could emit. The cap does not degrade the result, it
invalidates it.

SPA refusal: on a client-rendered site the links are not in the HTML and
every page reads as orphaned. That is catastrophic, so an empty graph returns
degraded/no_links_in_html instead of a full false-positive list. No JS
rendering by design — that is the R1/R2 arbitration, not something to smuggle
in here.

Verified against BOTH live sites and against a planted failure, because two
clean results are not evidence a detector detects:
- native PHP: 24 pages, 335 links, depth 2, 0 orphans
- Astro: 86 pages, 2015 links, depth 2, 0 orphans
- fixture with a planted orphan + a 4-click chain: both found. Filters proven
  on real shapes seen live — /css/main.css?v=1778157313, #anchors, mailto:,
  tel:, external hosts, .png. /b/ in markup vs /b in sitemap unify to one node
  rather than a phantom orphan pair.

Fixed a flaw in my own mock while writing that test: a single page.html
fixture cannot express a GRAPH (every node gets identical links), so the mock
is now pages.json = {url: html}.

Verified: seo-data 122 -> 136 pass, 0 fail; full suite green; shellcheck +
py_compile clean.
2026-07-17 12:32:42 +02:00
Bastien Chanot dca977bb27 fix(seo-data,seo): backtest on a second, native site — two real bugs
Everything on this branch was grounded on ONE Astro repo. A native PHP site
(lavageangels356.fr) broke two things that looked fine there.

BUG 1 — sitemap counted images as pages. _locs matched
`el.tag.endswith("}loc")`, and <image:loc> from Google's image-sitemap
namespace ALSO ends with '}loc'. Astro's sitemap has no image extension, so
this was invisible. The native site's does: 24 <url> + 3 <image:loc> came back
as count=27. The COVERAGE denominator was 12.5% too high and img/logo.png was
about to be sampled and audited as a page.
Fixed with two locks: walk the DIRECT children of each <url>/<sitemap> instead
of root.iter() (which alone excludes <image:image><image:loc>), and test the
sitemaps.org namespace explicitly. Regression fixture carries the image
extension; the old endswith code returns 9 URLs against it, the new one 7 with
zero images.
Verified both sites: native 27 -> 24, zero images; Astro unchanged at 86.

BUG 2 — the C1c family heuristic was tuned to one URL layout. "First path
segment" works for NESTED city pages (/creation-site-internet/essonne-91/ →
25 pages, 1 family) and FAILS for FLAT ones (/lavage-auto-pomponne,
/lavage-auto-torcy → 8 pages, 8 singletons). Consequence: C1c's rule "sample
>=3 from the largest family" would have targeted /services (5) and missed the
8 city pages entirely — the exact doorway-page risk the 30/70 rule exists to
catch.
Family is now "shared parent path OR shared slug prefix (>=3 URLs sharing 2+
hyphen tokens)", with both real layouts as the worked examples, plus a
sanity-check: a sitemap yielding almost as many families as URLs has defeated
the heuristic, not proved the site has no templates. Fixed in seo-analyzer and
in the geo pointer that referenced it.

Backtest results on the native site for everything else: url-guard accepts the
domain; source-scope excludes only .git (no dist/build/out exists — the
exclusions are correctly no-ops, and cache/ holds only .htaccess+.gitignore so
it is rightly untouched); the sameAs check runs and finds zero (a real GEO gap
for that site, not a tool bug); links are present in the served HTML (PHP is
SSR), so C3 is feasible there.

Verified: seo-data 119 -> 122 pass, 0 fail; full suite green; py_compile clean.
2026-07-17 12:28:36 +02:00
Bastien Chanot 3a15643c2c feat(seo-data): C2 — cannibalisation from Google's own data, one param away
The inventory called this "no duplicate-content / cannibalisation detection".
Splitting that into its two halves shows one is free and the other is a trap.

CANNIBALISATION — free, and the data was already reachable. Search Analytics
has always accepted several dimensions at once ("no limit to the number of
dimensions that you can group by"); this engine only ever sent
`"dimensions": [dim]` and _norm_queries only ever read `keys[0]`. So
query+page — the pairing that exposes the conflict — was one parameter away
and nobody asked. Same shape of win as W1.

  fetch.sh cannibal → queries where 2+ of OUR pages compete, ranked by total
  impressions, strongest page first inside each. Same auth, same quota family,
  no new scope. `capped` reports a full row window rather than presenting a
  truncated list as exhaustive — same rule as COVERAGE and the sitemap caps.

Grouping happens in the engine, deterministically: asking an LLM to group
1000 rows by query is arithmetic it should never be handed.

Backward compatible: rows gained `keys` (the list the API actually returns);
`key` stays as keys[0], so the single-dim quick-wins consumer is untouched.
A test pins both.

30/70 DUPLICATION — deliberately NOT built, and this is the honest half.
Measuring it needs main-content extraction (strip nav/header/footer). Without
that, comparing two same-template pages returns ~95% similar for every site —
a confident false positive, which is exactly the failure class the rest of
this branch exists to remove. It stays an explicit LLM judgement over the >=3
same-family pages C1c now samples for it, labelled as judgement, never quoting
a similarity percentage nobody computed. A wrong number would be worse than
the current honest gap.

The two must not be merged in the report either: cannibalisation is a SERP
fact Google measured; 30/70 is a content question. The spec now says so.

Verified: fixture with 3 pages on one query, 2 on another, 1 on a third →
2 conflicts, correct ranking, single-page query excluded; live dispatch
degrades cleanly with no account; seo-data 110 -> 119 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 11:55:00 +02:00
Bastien Chanot 04ccc5ad9b fix(seo,geo): C1c — split COVERAGE source/live; the 30/70 rule needed the
opposite sample

I5 made COVERAGE mandatory and told the agent to "sample by risk, one per
template". Grounding that against a real Astro site shows the rule is half
wrong, and that one number was hiding two.

86 URLs collapse into 8 families; 75 of them (87%) come from 3 dynamic
[dept] templates. So the same 12 sampled pages are simultaneously 14% LIVE
coverage and ~100% SOURCE coverage. Reporting only 14% understates the audit;
reporting only 100% oversells it. Both lines now, in both agents — and they
bound different findings, so they must not be averaged:
  SOURCE bounds CODE (one template renders its whole family: a missing
  canonical in [dept]/index.astro breaks all 25 identically).
  LIVE bounds CONTENT (title wording, thin pages, duplication — written per
  page, so a template says nothing about its 25 instances).

The sharper half: "one per template" is CORRECT for code and WRONG for the
30/70 rule, which this same spec mandates at :954 (city pages: 30% shared,
70% unique). You cannot tell whether 25 dept pages are 70% unique by reading
one of them. The spec was mandating a check its own sampling made
structurally impossible. Sampling is now keyed to the finding class: 1 per
family for code, >=3 from the LARGEST family for duplication, spread for
per-page content.

Families come from the first path segment of the sitemap URLs (C1b) — a good
enough proxy for "same template" that needs no framework routing knowledge,
verified against the real distribution.

geo gets the same split, cut differently: JSON-LD lives in a shared layout so
SOURCE bounds Schema.org, while Definition Lead / TL;DR are written per page
so LIVE bounds Content Shape. Site-wide axes (crawler policy, llms.txt) stay
unbounded — single files, fully read.

Verified: full suite green. Caught and fixed one self-inflicted contradiction
before commit — geo's prose demanded both coverages while its output block
still had a single line.
2026-07-17 11:43:48 +02:00
Bastien Chanot 2de58faa38 feat(seo-data): C1b — sitemap verb, the denominator COVERAGE never had
I5 made a COVERAGE line mandatory in STEP 9 and told the agent to "count the
URLs in sitemap.xml" without giving it a command. STEP 4 only ever did
`curl … | head -50` — a preview, not a count. This closes that.

fetch.sh sitemap --url … → {count, urls[], index, dropped}. Stdlib only
(urllib + xml.etree + gzip): no auth, no Google, no venv, so it runs wherever
the mock/degrade paths run. Follows <sitemapindex> one level, dedupes, strips
whitespace, handles .xml.gz. Every cap REPORTS what it cut (children_skipped,
truncated) rather than truncating silently — same rule as COVERAGE itself.

PLAN CORRECTION: the proposal said the verb would "validate each URL via the
H1 guard". Wrong. urllib fetches these, so nothing here reaches a shell and
there is no injection surface to guard. The guard belongs at the point of
use, where seo-analyzer interpolates a URL into curl — which is the contract
the sameAs check already established. A second copy of url-guard here would
only drift from the first. The module carries a garbage filter, named as such.

SECURITY: the security-guidance hook asked for defusedxml. Taken seriously,
not obeyed — it would drag a venv into a module whose whole point is being
stdlib-only. Split the threat instead: xml.etree does NOT expand external
entities (XXE is not the vector), but it IS billion-laughs-vulnerable, and
the 20 MB read ceiling bounds the input, not the expansion. A sitemap NEVER
has a DTD — sitemaps.org is <?xml?> then <urlset xmlns=> — so any
doctype/entity is refused BEFORE parsing, with its own reason
(unsafe_xml_dtd, distinct from parse_failed: it is a finding, not a glitch).
Refusing the construct beats depending on parser internals. Fixture is a real
billion-laughs payload.

Verified against the live target, not just fixtures: zenquality's sitemap
returns count=86, dropped=0, matching `grep -c '<loc>'` on the raw XML
exactly. Dead URL → {"status":"degraded","reason":"fetch_failed"}, exit 0.
seo-data 95 -> 110 pass, 0 fail; full suite green; shellcheck + py_compile
clean.

Note: no config-edit sentinel was needed after all — config-protection guards
lib/tests, not lib/seo-data. I posted one, found it uncommitted-and-unconsumed
afterwards, and removed it rather than leave an open one-shot gate lying
around. Worth knowing: seo-data.test.sh is 110 assertions and is NOT covered
by that hook, while lib/tests/*.test.sh is.
2026-07-17 11:28:00 +02:00
Bastien Chanot 8dcdc661ce fix(seo,geo): C1a — find sees build output, grep does not; the two disagreed
Surfaced by dogfooding on a real Astro repo instead of reading the spec.

Claude Code installs a shell function routing `grep` to ugrep with
`--ignore-files`, so grep honours .gitignore and never descends into a
gitignored dist/. `find` honours nothing. seo-analyzer uses both.

Measured on zenquality (Astro, dist/ gitignored but built locally),
seo-analyzer.md:497 returned 92 images, 45 of them under dist/ — every asset
listed twice, source and generated copy, byte-identical. Two live
consequences:
- "top 20 by size" was ~10 real images dressed as 20.
- Batch C (`cwebp -q 80 <img> -o <img>.webp`) could target dist/og-image.png;
  the .webp lands in dist/ and the `npm run build` the dispatcher runs to
  VERIFY the fix erases it. Fix lands, verification passes, nothing survives,
  report says applied.

lib/source-scope.sh separates source from build output, framework-aware.
`public/` is deliberately NOT excluded by default: it is Astro/Vite/Next
SOURCE and holds favicon.ico, apple-touch-icon.png and robots.txt — the very
files STEP 4 curls. It is build output only for Hugo/Gatsby, detected from
config (legacy config.toml alone is ambiguous, so it needs archetypes/ too).
Blanket-excluding it would blind the audit to its own resource checks.

findargs emits one token per line and MUST be consumed via a quoted array.
I shipped a flat-string version first and the dogfood caught it: the shell
globs */dist/* against the CWD and passes the matches to find as search
paths, which turned 90 hits into 135 and kept every dist/ file. Both the
header and a functional test now pin that.

Also: no bundle item may target build output, in either agent. That is the
real safety net — even if some future find leaks a dist path, the fix cannot
land there. Fix the source that generates the artifact; if the source cannot
be found, that is a finding, not a reason to patch the artifact.

SCOPE CORRECTION: the proposal claimed grep was auditing 86 generated files
instead of 9 templates. That was FALSE — the ugrep shim already skips them.
Killed my own premise before coding it; the real bug is narrower and lives in
find only. Do NOT "fix" the grep lines to match: they are already correct,
and adding these exclusions there would drop public/.

Verified: 90 -> 45 images on the real repo, 0 dist survivors, public/
preserved (favicon.ico still visible); 34 new assertions PASS / 0 FAIL; full
suite green; shellcheck clean. Test file addition used the documented
one-shot config-edit sentinel.
2026-07-17 09:50:57 +02:00
Bastien Chanot 7d6aa09faf feat(lib): H1 — url-guard, shell-injection + local-target refusal before curl
Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is
typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+,
geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the
threat model completely: URLs then come from the TARGET'S OWN SERVER, so a
remote file's bytes reach a shell.

The severe hazard is injection, not SSRF. Those curls quote with ", inside
which $ and backtick still execute, and ~/.claude/.env holds
GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A <loc> of
`https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The
test suite asserts exactly that payload is refused.

Code, not prose: a markdown instruction does not stop an injection. Mirrors
the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C
locale, POSIX case: newline-proof, locale-independent, no grep pitfall.
Allowlist over denylist per CLAUDE.md.

Covers: shell metacharacters; scheme (http/https only — no file:, gopher:);
literal loopback/private/link-local/metadata/.local; userinfo authority
confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com).

NOT covered, stated in the header rather than left silent: DNS-level SSRF. A
public hostname resolving to a private address passes. Closing it needs
resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU
window. Proportionate to the threat model — this runs on a workstation
auditing the operator's own client sites.

Wired at all three entry points: both agents' STEP 4 domain assignment, and
the W3 sameAs loop (whose URLs come from the audited repo, not the operator).
Refused sameAs rows report as REFUSED rather than vanish — neither dead nor
live, and an unguardable sameAs is itself a finding.

Note: writing the test file tripped the config-protection hook (test suite is
a guarded quality-gate). Used the documented one-shot sentinel with a reason
rather than working around the gate; it was consumed as designed.

Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite
green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the
health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded
against the real zenquality.fr domain (accepted) and the real exfil payload
(refused, exit 2).
2026-07-17 09:25:34 +02:00
Bastien Chanot a6d423b940 feat(seo-data): W1 — surface rich_results, the data inspect already threw away
google_seo.py:129 read only indexStatusResult out of the URL Inspection
response and discarded the rest. richResultsResult was already on the wire:
same call, same OAuth scope (webmasters.readonly), same quota. Google's own
structured-data verdict on the live indexed URL was being downloaded and
binned.

Plan correction: the TODO said "richresults verb". Wrong — a new verb means
a second POST to the same endpoint for a payload already received, on a
per-site quota, and nobody wants rich results without index status. Extended
inspect() instead; fetch.sh unchanged, no new verb, no new scope.

Design driven by the published schema, not by guesswork — two details I
would have got wrong:
- richResultsResult is OMITTED when Google detects none ("absent if none
  found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a
  caller cannot tell an absent key from a check that never ran. ABSENT means
  "none detected", never "invalid". The KeyError path is the real risk here,
  so it has its own fixture dir (fixtures-norich/) and its own tests.
- PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It
  never emits it now, and a test asserts the absence.

issues[] deduped (the same issueMessage repeats across every affected item),
errors/warnings count instances — scale from the counter, cause from the
message.

seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD
validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a
verified property, so its reach is the STEP 9 COVERAGE ratio, not the site.
Replacing a fake validator with a fake coverage promise would be no better.

This is what beats claude-seo: their README's "dual validator (Rich Results
Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of
their .py finds zero calls. This is Google's verdict, via auth already held.

Note: the new dedupe assertion trips SC2015 (A && B || C), same as the
pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for
house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh
shellcheck glob anyway.

Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end
and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean.
2026-07-16 20:51:30 +02:00
Bastien Chanot fe93b7945b feat(geo): W3 — implement the sameAs resolution check that the spec promised
entity-seo.md:148 says "sameAs pointing to dead profiles — validate each URL
resolves", and STEP 7 asks "does the target resolve and match?". Nothing
implemented it: zero curl against a sameAs anywhere in the repo. A dead
sameAs is worse than a missing one — it asserts an identity link that fails
on follow, in the exact graph AI engines walk to confirm who you are.

The naive version of this check is a false-positive generator, which is
presumably why it stayed unimplemented. Verified live rather than assumed:

  999  linkedin.com/company/anthropic   <- blocks non-browsers
  200  wikidata.org/wiki/Q108162414
  200  x.com/anthropicai
  404  <known-dead URL>                 <- correctly detected

So the check classifies by code, not by liveness guess: 404/410 = dead
(finding with direction), 401/403/429/999 = bot-blocked (inconclusive, NO
finding, never "dead"), 000/5xx = inconclusive. No G2/G6 item may remove a
sameAs on anything but 404/410 — same shape as the NAP direction rule: an
unreliable signal read confidently is worse than no signal.

Note: the spec draft asserted "X/Twitter and Instagram commonly 403" from
plausibility. The live test returned 200 for x.com and contradicted it —
corrected to classify by observed code, never by platform folklore. Third
unverified-plausible claim caught this session (I1, I6, here); the pattern
is exactly what these fixes exist to stop.

Verified: pipeline exercised end-to-end against real endpoints; make test
35 GREEN / 0 RED.
2026-07-16 20:39:17 +02:00
Bastien Chanot acd452b92f fix(seo): I8 — drop the phantom .claude/audits/external/ precondition
STEP 0 told the user to run `mkdir -p .claude/audits/external` themselves
before handing over an external report. Three things wrong with that:

- The skill runs dozens of bash commands but outsourced this one to a human.
- The timing was impossible: to "drop the export in" that directory the user
  needed it to already exist, so the instruction arrived after the moment it
  would have been useful.
- The directory is not needed at all. `:218` already reads "File path given
  → Read it" — any path works — and nothing in skills/ or agents/ ever
  writes to that path. Grep confirms it is referenced by exactly these two
  lines and known to nothing else: a convention the skill invented, asked
  the user to create, and never used.

Fix removes the precondition instead of automating it: give a path from
anywhere, the tidy location stays a suggestion.

Verified: make test 35 GREEN / 0 RED.
2026-07-16 20:36:03 +02:00
Bastien Chanot 9da1dec9e6 fix(geo): I6 — every stat was real and attached to the wrong claim
Audited each statistic in agents/resources/ against primary sources after
the VSI fiction (I2) showed WebSearch launders SEO-blog consensus.

The failure mode is not invention — it is plausible recombination, which is
what a model half-remembering a search result produces:

- "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)"
  — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate
  over the whole method set, domain-dependent. No per-technique figure
  exists.
- "Pages not updated quarterly are 3x more likely to lose AI citations
  (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more
  strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source
  supports a quarterly decay multiplier.
- "QAPage cited 58% more often than Article" — uncited. Nearest real number:
  AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources —
  wrong type, and its FAQPage figure (1.8%) points the opposite way to the
  claim it propped up. This one drove Tier 1 ranking.
- "62% of searches involve voice" — uncited; 62% circulates as smart-speaker
  ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a
  2014 Andrew Ng interview).

Corrected my own framing too: I claimed three times these stats "drive axis
weights". They do not — the weight tables carry no citations. They drive
Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule
pushed them into CLIENT reports as research-backed.

Fixes: recommendations kept on mechanism, fabricated numbers removed with
the incident documented inline so they are not re-added. Unverified stats
(48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED]
rather than asserted or deleted — I did not check them.

Structural, not just exhortation: resources/README.md now mandates
`<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>`. `measured:` is the field that catches this — all four
errors survive a source name; none survives stating the real measurement
next to the claim. WebSearch demoted from verification to crawler/tool-name
lookup only.

Verified: make test 35 GREEN / 0 RED.
2026-07-16 20:32:55 +02:00
Bastien Chanot e70e1d6c71 fix(seo): I4 — stop double-counting security headers; /harden owns them
Headers were scored three ways: seo-analyzer priced them into the Technical
axis at both depths (:619 FULL, :635 LOCAL), depth-matrix.md:29 said drop
them, and /harden re-audits them 0-100 against three external validators.
The dedup rule and the agent spec contradicted each other; the agent won by
default, so the same finding moved two scores in two reports.

Arbitrated (user): /harden keeps them, /seo drops them. That confirms the
rule that already existed — seo-analyzer was the violator.

Constraint: /harden REUSES seo-analyzer, so the capability cannot be
deleted, only scoped. Reading is not scoring:
- Technical axis definitions no longer name security headers.
- STEP 4 still curls them — needed for X-Robots-Tag, canonical/redirect
  coherence, and the §14 observed-list — but they earn no points under /seo.
- Dispatched from /harden: unchanged, headers ARE the job (verified: its
  scope spec untouched, 16 header references intact).

Carve-out: X-Robots-Tag stays in /seo under indexability. It is an indexing
directive wearing a header's clothes — `noindex` there deindexes as surely
as a meta robots tag. That is what depth-matrix.md:29 means by "unless it
directly affects indexability"; the security headers do not.

Drop is not silence: mandatory §14 line on FULL naming what was observed
live plus a "run /harden <url>" pointer. A user who never runs /harden must
not read a clean Technical score as clean headers — same principle as the
mandatory COVERAGE line (I5).

Verified: make test 35 GREEN / 0 RED.
2026-07-16 16:56:32 +02:00
Bastien Chanot 64f175f01d fix(seo,geo): I5 — disclose sampling coverage instead of implying an audit
Both agents sample (seo-analyzer.md:403 "sample 5-15 key pages",
geo-analyzer.md:460 "Sample 5-10 key pages") and neither states it. The
report says "audit". On a 500-page site a 12-page sample is 2.4%, and the
reader cannot know that unless it is printed. /client-handover gates on
these scores.

The denominator was already within reach: STEP 4 fetches sitemap.xml. Count
its URLs and the coverage ratio is free — same data C1 will use for
sitemap-driven crawl later.

- Mandatory COVERAGE line in both scoring blocks: N of M sitemap URLs (P%),
  or "total UNKNOWN" when no sitemap. Never omitted, never rounded up.
- seo: <25% coverage repeats in §0 as a major alert. Sample by risk (one
  page per template + GSC position 4-10 quick wins), name skipped templates
  — an un-sampled template is an un-audited template.
- geo: scoped honestly rather than blanket — COVERAGE bounds the per-page
  axes (Content Shape, page-level Schema.org) but NOT the site-wide ones
  (AI Crawlers Policy, llms.txt are single files, fully read). One ratio
  should not discredit axes it does not govern.

Verified: make test 35 GREEN / 0 RED.
2026-07-16 16:50:19 +02:00
Bastien Chanot 4ea2fb8c37 fix(seo): I2 — remove VSI, an SEO-blog fiction, from CWV thresholds
seo-analyzer.md:278 listed "VSI (Visual Stability Index) — new 2026 signal,
Google Core Web Vitals 2.0" as a threshold, stated as fact, no hedge, in
client-facing audits. It does not exist.

Verified against two primary sources:
- developer.chrome.com/docs/crux/api — complete metric list carries no
  visual_stability_index. Blogs claimed "Google is actively collecting VSI
  through CrUX": flatly false.
- web.dev/articles/vitals — three stable CWV (LCP, INP, CLS). No VSI, no
  "Core Web Vitals 2.0". Thresholds change with prior notice on an annual
  cadence.

Ten SEO blogs cross-cited each other into an apparent consensus. WebSearch
returns that consensus, which is why the resources README rule "agents MUST
cross-check via WebSearch on FULL" did not catch it — that mitigation
launders blog misinformation into apparent verification.

Fix removes the metric and states the sourcing rule where a future rumour
would land: primary sources only (web.dev / Chromium blog / CrUX API list,
the last being decisive — a metric CrUX cannot return is one we cannot
score). Incident documented inline so it is not re-added.

Verified: make test 35 GREEN / 0 RED.
2026-07-16 16:46:52 +02:00
Bastien Chanot 9cd7b51bb8 fix(seo,geo): dogfood on zenquality.fr — two process anomalies
Surfaced by pointing /harden at zenquality.fr from the claude-config CWD.

A1 — no CWD/target coherence guard (systemic: /seo, /geo, /harden all
lack it; grep confirms). A URL is supplied, the agent greps whatever CWD
it landed in, nobody checks they are the same site. Demonstrated live:
from claude-config, /harden would curl zenquality.fr while grepping
claude-config, then score "Config hardening" on a codebase that is not the
site. The live half looks right, the code half is fiction, and the report
reads as authoritative.
Fixed in both agents' STEP 2 rather than the 3 dispatchers: the agent does
the grepping, so the guard binds whoever calls — same principle as I3.
/harden inherits it free.

A2 — seo-analyzer had zero origin-vs-edge awareness while geo-analyzer
has the full CDN/WAF-override check (geo-analyzer.md:246-261). seo-analyzer
does the infra detection AND is reused by /harden for its whole
config-hardening axis (20/100). On zenquality — Apache origin behind a
Scaleway nginx front — repo .htaccess + `server: nginx` invites the wrong
call "nginx serves this, .htaccess is dead". I made that exact inference
myself before reading the file. Rule added at STEP 2 infra detection:
`server:` names the edge, not the origin; live-but-not-in-repo = "set
upstream", never "missing".

Verified: live probe of zenquality.fr (read-only, nothing written to the
client repo); make test 35 GREEN / 0 RED.
2026-07-16 16:25:16 +02:00