From fe93b7945bffe1372f12fe476366e3e37ba61ae3 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:39:17 +0200 Subject: [PATCH 01/14] =?UTF-8?q?feat(geo):=20W3=20=E2=80=94=20implement?= =?UTF-8?q?=20the=20sameAs=20resolution=20check=20that=20the=20spec=20prom?= =?UTF-8?q?ised?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit entity-seo.md:148 says "sameAs pointing to dead profiles — validate each URL resolves", and STEP 7 asks "does the target resolve and match?". Nothing implemented it: zero curl against a sameAs anywhere in the repo. A dead sameAs is worse than a missing one — it asserts an identity link that fails on follow, in the exact graph AI engines walk to confirm who you are. The naive version of this check is a false-positive generator, which is presumably why it stayed unimplemented. Verified live rather than assumed: 999 linkedin.com/company/anthropic <- blocks non-browsers 200 wikidata.org/wiki/Q108162414 200 x.com/anthropicai 404 <- correctly detected So the check classifies by code, not by liveness guess: 404/410 = dead (finding with direction), 401/403/429/999 = bot-blocked (inconclusive, NO finding, never "dead"), 000/5xx = inconclusive. No G2/G6 item may remove a sameAs on anything but 404/410 — same shape as the NAP direction rule: an unreliable signal read confidently is worse than no signal. Note: the spec draft asserted "X/Twitter and Instagram commonly 403" from plausibility. The live test returned 200 for x.com and contradicted it — corrected to classify by observed code, never by platform folklore. Third unverified-plausible claim caught this session (I1, I6, here); the pattern is exactly what these fixes exist to stop. Verified: pipeline exercised end-to-end against real endpoints; make test 35 GREEN / 0 RED. --- agents/geo-analyzer.md | 41 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 41 insertions(+) diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index b3034db..38a0cab 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -431,6 +431,47 @@ Record what exists. For each: - Does `sameAs` on the site point to it? - If yes, does the target resolve and match? +### sameAs resolution `[FULL only]` + +`entity-seo.md:148` says "validate each URL resolves" and nothing did. +A `sameAs` pointing at a dead profile is worse than a missing one: it +asserts an identity link that fails on follow, in the exact graph AI +engines walk to confirm who you are. + +```bash +grep -rhoE '"sameAs"[^]]*\]' \ + --include="*.html" --include="*.astro" --include="*.tsx" --include="*.jsx" \ + --include="*.vue" --include="*.svelte" --include="*.php" --include="*.json" \ + . 2>/dev/null \ + | grep -oE 'https?://[^"]+' | sort -u | while read -r U; do + printf '%s %s\n' \ + "$(curl -sIL -o /dev/null -w '%{http_code}' --max-time 10 "$U" 2>/dev/null || echo 000)" \ + "$U" + done +``` + +**Read the codes honestly — a block is not a death.** Some platforms refuse +non-browser clients: LinkedIn answers `999` (verified 2026-07-16 against a +live company page). A naive check calls that dead and the bundle deletes a +live link — the most valuable node in the graph, since LinkedIn is the +identity anchor for most B2B entities. + +Do NOT assume which platforms block: the same 2026-07-16 check found +`x.com` returning `200`, contradicting the "Twitter always 403" folklore. +Test the code you actually got; classify by code, never by platform +reputation. + +| Code | Verdict | Action | +|---|---|---| +| 2xx / 3xx | alive | none | +| **404 / 410** | **genuinely dead** | finding WITH direction — fix or remove | +| 401 / 403 / 429 / 999 | bot-blocked | **inconclusive — no finding.** Report as unverified, never as dead | +| 000 (DNS/timeout) / 5xx | inconclusive | retry once, then unverified | + +No G2/G6 item may remove a `sameAs` on anything but 404/410. Same rule as +the NAP direction rule: an unreliable signal read confidently is worse than +no signal. Unverified entries → §14, naming the platform and the code. + ### Google Knowledge Panel `[FULL only]` ``` From a6d423b940740055f59de4e226629a70489576f7 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:51:30 +0200 Subject: [PATCH 02/14] =?UTF-8?q?feat(seo-data):=20W1=20=E2=80=94=20surfac?= =?UTF-8?q?e=20rich=5Fresults,=20the=20data=20inspect=20already=20threw=20?= =?UTF-8?q?away?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit google_seo.py:129 read only indexStatusResult out of the URL Inspection response and discarded the rest. richResultsResult was already on the wire: same call, same OAuth scope (webmasters.readonly), same quota. Google's own structured-data verdict on the live indexed URL was being downloaded and binned. Plan correction: the TODO said "richresults verb". Wrong — a new verb means a second POST to the same endpoint for a payload already received, on a per-site quota, and nobody wants rich results without index status. Extended inspect() instead; fetch.sh unchanged, no new verb, no new scope. Design driven by the published schema, not by guesswork — two details I would have got wrong: - richResultsResult is OMITTED when Google detects none ("absent if none found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a caller cannot tell an absent key from a check that never ran. ABSENT means "none detected", never "invalid". The KeyError path is the real risk here, so it has its own fixture dir (fixtures-norich/) and its own tests. - PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It never emits it now, and a test asserts the absence. issues[] deduped (the same issueMessage repeats across every affected item), errors/warnings count instances — scale from the counter, cause from the message. seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a verified property, so its reach is the STEP 9 COVERAGE ratio, not the site. Replacing a fake validator with a fake coverage promise would be no better. This is what beats claude-seo: their README's "dual validator (Rich Results Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of their .py finds zero calls. This is Google's verdict, via auth already held. Note: the new dedupe assertion trips SC2015 (A && B || C), same as the pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh shellcheck glob anyway. Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean. --- agents/seo-analyzer.md | 29 +++++++++++++ lib/seo-data/README.md | 17 +++++++- lib/seo-data/fixtures-norich/gsc_inspect.json | 2 + lib/seo-data/fixtures/gsc_inspect.json | 13 +++++- lib/seo-data/google_seo.py | 41 ++++++++++++++++++- lib/seo-data/seo-data.test.sh | 18 ++++++++ 6 files changed, 115 insertions(+), 5 deletions(-) create mode 100644 lib/seo-data/fixtures-norich/gsc_inspect.json diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 190a865..8c7070d 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -339,6 +339,35 @@ and 10 AND high impressions (candidates to push onto page 1 with a title/meta/content tweak). Report index coverage from `inspect`. All emitted into SEO.md §2 (technical) and §8 (quick wins). +**`inspect` also returns `rich_results` — Google's own structured-data +verdict on the live indexed URL.** It rides the same response (no extra +call, no extra quota). This is the only programmatic JSON-LD validation in +the system; everything else about schema is read by eye. + +``` +rich_results.verdict : PASS | FAIL | NEUTRAL | VERDICT_UNSPECIFIED | ABSENT +rich_results.types[] : {type, items, errors, warnings, issues[]} +``` + +- `FAIL` + a type carrying `errors > 0` → that type **cannot show as a rich + result**. Bundle item, cite the `issues[]` message verbatim — it is + Google's wording, not ours, and geo-analyzer owns the JSON-LD fix + (CROSS-AGENT NOTE). +- `warnings` → recommended fields missing. Report, do not gate on them. +- **`ABSENT` means Google detected no rich results on this URL** — the key + is omitted upstream when nothing is found. It is NOT an error and NOT + proof the markup is broken: a page with no structured data reads the same + as one whose markup Google never parsed. Say "none detected", never + "invalid". +- `ABSENT` while the repo clearly ships JSON-LD → real finding: the markup + is not reaching Google (SPA-rendered, blocked, or malformed). Cross-check + before claiming it. + +**Bound this honestly.** `index:inspect` is per-URL, quota'd, and works only +on a GSC-verified property. It validates the URLs you sampled — not the +site. Its reach is the STEP 9 COVERAGE ratio, and §14 must say so rather +than let one PASS imply site-wide valid markup. + If `status=degraded` → note it in §2 and emit the §11 user action "Connecter GSC: `make seo-connect`". diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index b730beb..c21b097 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -80,9 +80,24 @@ fetch.sh queries --account client-a --property sc-domain:ex.com [--days 90] [--d → {"status":"degraded","reason":"no_credentials"|"token_revoked"|"network_error"|"rate_limited"} fetch.sh inspect --account client-a --property … --url https://ex.com/page - → {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…"} + → {"status":"ok","source":"gsc","indexed":true,"coverage":"…","last_crawl":"…", + "rich_results":{"verdict":"PASS|FAIL|NEUTRAL|VERDICT_UNSPECIFIED|ABSENT", + "types":[{"type":"FAQ","items":2,"errors":2,"warnings":1, + "issues":["Missing field 'acceptedAnswer'"]}]}} → {"status":"degraded","reason":"…"} + rich_results rides the SAME URL-Inspection response — Google already sends + it, `inspect` used to discard it. No extra call, quota or OAuth scope. + It is the only programmatic structured-data validation in the system. + • verdict PARTIAL is never emitted — the API reserves it as unused. + • verdict ABSENT is SYNTHETIC (not a Google enum): the API omits + richResultsResult entirely when it detects no rich results. Surfaced + as a value rather than a missing key, because a caller cannot tell an + absent key apart from a check that never ran. ABSENT = "none + detected", never "invalid". + • errors/warnings count issue INSTANCES; issues[] is deduped — the same + issueMessage repeats across every affected item. + fetch.sh forget --label client-a → {"status":"ok","removed":true|false} # false = label wasn't in the store diff --git a/lib/seo-data/fixtures-norich/gsc_inspect.json b/lib/seo-data/fixtures-norich/gsc_inspect.json new file mode 100644 index 0000000..325cacd --- /dev/null +++ b/lib/seo-data/fixtures-norich/gsc_inspect.json @@ -0,0 +1,2 @@ +{"inspectionResult":{"indexStatusResult":{ + "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}} diff --git a/lib/seo-data/fixtures/gsc_inspect.json b/lib/seo-data/fixtures/gsc_inspect.json index 325cacd..bb5de1f 100644 --- a/lib/seo-data/fixtures/gsc_inspect.json +++ b/lib/seo-data/fixtures/gsc_inspect.json @@ -1,2 +1,11 @@ -{"inspectionResult":{"indexStatusResult":{ - "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}}} +{"inspectionResult":{ + "indexStatusResult":{ + "verdict":"PASS","coverageState":"Submitted and indexed","lastCrawlTime":"2026-07-01T10:00:00Z"}, + "richResultsResult":{"verdict":"FAIL","detectedItems":[ + {"richResultType":"Breadcrumbs","items":[{"name":"Unnamed item","issues":[]}]}, + {"richResultType":"FAQ","items":[ + {"name":"Q1","issues":[ + {"issueMessage":"Missing field 'acceptedAnswer'","severity":"ERROR"}]}, + {"name":"Q2","issues":[ + {"issueMessage":"Missing field 'acceptedAnswer'","severity":"ERROR"}, + {"issueMessage":"Unspecified image","severity":"WARNING"}]}]}]}}} diff --git a/lib/seo-data/google_seo.py b/lib/seo-data/google_seo.py index d73277d..a544c6f 100644 --- a/lib/seo-data/google_seo.py +++ b/lib/seo-data/google_seo.py @@ -114,6 +114,39 @@ def queries(store_path, account, property, days=90, dim="query"): raw = r.json() return _norm_queries(raw, dim) +def _rollup_issues(items): + """Count issue instances by severity; dedupe messages (they repeat per item).""" + errors = warnings = 0 + msgs = [] + for item in items: + for iss in item.get("issues", []): + sev = iss.get("severity") + if sev == "ERROR": + errors += 1 + elif sev == "WARNING": + warnings += 1 + msg = iss.get("issueMessage") + if msg and msg not in msgs: + msgs.append(msg) + return errors, warnings, msgs + +def _norm_rich(ir): + """richResultsResult → verdict + per-type rollup. Google OMITS the key when + it detects no rich results, so absence is data, not an error: surfaced as the + synthetic verdict ABSENT (not a Google enum) rather than a missing key, which + a caller cannot tell apart from a check that never ran. PARTIAL is never + emitted — the API reserves it as unused.""" + rr = ir.get("richResultsResult") + if rr is None: + return {"verdict": "ABSENT", "types": []} + types = [] + for det in rr.get("detectedItems", []): + errors, warnings, msgs = _rollup_issues(det.get("items", [])) + types.append({"type": det.get("richResultType"), + "items": len(det.get("items", [])), + "errors": errors, "warnings": warnings, "issues": msgs}) + return {"verdict": rr.get("verdict"), "types": types} + def inspect(store_path, account, property, url): raw = _mock("gsc_inspect.json") if raw is None: @@ -126,11 +159,15 @@ def inspect(store_path, account, property, url): return {"status": "degraded", "reason": "rate_limited"} r.raise_for_status() raw = r.json() - isr = raw["inspectionResult"]["indexStatusResult"] + ir = raw["inspectionResult"] + isr = ir["indexStatusResult"] + # rich_results rides the SAME response — Google already sent it and this + # function used to discard it. No extra call, no extra quota, no new scope. return {"status": "ok", "source": "gsc", "indexed": isr.get("verdict") == "PASS", "coverage": isr.get("coverageState"), - "last_crawl": isr.get("lastCrawlTime")} + "last_crawl": isr.get("lastCrawlTime"), + "rich_results": _norm_rich(ir)} def _cli(): try: diff --git a/lib/seo-data/seo-data.test.sh b/lib/seo-data/seo-data.test.sh index 16e8a81..751a8e0 100644 --- a/lib/seo-data/seo-data.test.sh +++ b/lib/seo-data/seo-data.test.sh @@ -57,6 +57,24 @@ has "queries position field" "$Q" '"position": 6.3' I="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/google_seo.py" inspect \ --store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)" has "inspect indexed true" "$I" '"indexed": true' +# rich_results rides the same URL-Inspection response (no extra call/quota) +has "rich verdict surfaced" "$I" '"verdict": "FAIL"' +has "rich type breadcrumbs" "$I" '"type": "Breadcrumbs"' +has "rich type faq" "$I" '"type": "FAQ"' +has "rich counts error severity" "$I" '"errors": 2' +has "rich counts warn severity" "$I" '"warnings": 1' +has "rich keeps issue message" "$I" "Missing field 'acceptedAnswer'" +# same issueMessage repeats across items — the rollup must collapse it to one +NMSG="$(printf '%s' "$I" | grep -cF "Missing field 'acceptedAnswer'")" +[ "$NMSG" = "1" ] && ok "rich dedupes issue messages" \ + || no "rich dedupes issue messages" "got $NMSG occurrences" +# Google OMITS richResultsResult when it detects none — absence is data, and +# must not KeyError nor vanish into a missing key +NR="$(SEO_DATA_MOCK_DIR="$SD/fixtures-norich" python3 "$SD/google_seo.py" inspect \ + --store "$S2" --account client-a --property sc-domain:ex.com --url https://ex.com/x)" +has "no-rich → synthetic ABSENT" "$NR" '"verdict": "ABSENT"' +has "no-rich keeps index status" "$NR" '"indexed": true' +hasnt "no-rich emits no PARTIAL" "$NR" 'PARTIAL' DEG="$(env -u SEO_DATA_MOCK_DIR python3 "$SD/google_seo.py" queries \ --store "$TMP2/none.json" --account nobody --property sc-domain:ex.com)" has "gsc degrades w/o creds" "$DEG" '"status": "degraded"' From 7d6aa09faf21e25516d415f670a02e1c18d4cfed Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 09:25:34 +0200 Subject: [PATCH 03/14] =?UTF-8?q?feat(lib):=20H1=20=E2=80=94=20url-guard,?= =?UTF-8?q?=20shell-injection=20+=20local-target=20refusal=20before=20curl?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+, geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the threat model completely: URLs then come from the TARGET'S OWN SERVER, so a remote file's bytes reach a shell. The severe hazard is injection, not SSRF. Those curls quote with ", inside which $ and backtick still execute, and ~/.claude/.env holds GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A of `https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The test suite asserts exactly that payload is refused. Code, not prose: a markdown instruction does not stop an injection. Mirrors the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C locale, POSIX case: newline-proof, locale-independent, no grep pitfall. Allowlist over denylist per CLAUDE.md. Covers: shell metacharacters; scheme (http/https only — no file:, gopher:); literal loopback/private/link-local/metadata/.local; userinfo authority confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com). NOT covered, stated in the header rather than left silent: DNS-level SSRF. A public hostname resolving to a private address passes. Closing it needs resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU window. Proportionate to the threat model — this runs on a workstation auditing the operator's own client sites. Wired at all three entry points: both agents' STEP 4 domain assignment, and the W3 sameAs loop (whose URLs come from the audited repo, not the operator). Refused sameAs rows report as REFUSED rather than vanish — neither dead nor live, and an unguardable sameAs is itself a finding. Note: writing the test file tripped the config-protection hook (test suite is a guarded quality-gate). Used the documented one-shot sentinel with a reason rather than working around the gate; it was consumed as designed. Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded against the real zenquality.fr domain (accepted) and the real exfil payload (refused, exit 2). --- .claude/tasks/TODO.md | 50 ++++++++++++++++++++++- agents/geo-analyzer.md | 20 ++++++++- agents/seo-analyzer.md | 9 ++++- lib/tests/url-guard.test.sh | 74 +++++++++++++++++++++++++++++++++ lib/url-guard.sh | 81 +++++++++++++++++++++++++++++++++++++ 5 files changed, 230 insertions(+), 4 deletions(-) create mode 100644 lib/tests/url-guard.test.sh create mode 100644 lib/url-guard.sh diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 2f78ca7..d60a28b 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,6 +1,54 @@ # TODO -## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage) +## 2026-07-17 — STATUS seo/geo parity (branch bugfix/seo-geo-integrity, 10 commits, UNMERGED) +PHASE 1 — integrity: **DONE 7/7**. I3 8b0c98c · I1 57c67f2 · I2 4ea2fb8 · +I5 64f175f · I4 e70e1d6 · I6 9da1dec · I8 acd452b. Plus 9cd7b51 (A1+A2, two +process anomalies surfaced by dogfooding /harden at zenquality.fr from the +wrong CWD). +PHASE 2 — free wins: W3 fe93b79 · W1 a6d423b · **W2 DEFERRED** (see below). +NEXT: H1 (SSRF/injection guard) → C1 (sitemap crawl). Human merge gate: all +10 commits await review; nothing merged to develop. + +### Plan corrections made while executing (the plan was wrong 4×) +- **B3 KILLED** — GSC Links API does not exist. Verified against the API + reference: Search Console v1 exposes exactly Search Analytics, Sitemaps, + Sites, URL Inspection. A subagent hallucinated it; I doubted it in the + plan and the doubt was right. Common Crawl is the ONLY free backlink + source → the 70/100 cap is mandatory, not optional. +- **I1 was an over-correction** — "Off-page has ZERO data" was overstated + (relayed from a subagent, unverified). Brand mentions ARE gathered + (STEP 6). Narrowed the axis definition instead of N/A-ing it; weights + untouched to avoid churning historical scores twice. +- **I6 framing was wrong** — I claimed 3× that the stats "drive axis + weights". They do not; weight tables carry no citations. They drive Tier + recommendations and, worse, land in CLIENT reports via the "Cite sources" + rule. Reality was worse than my false version. +- **W1 was the wrong shape** — plan said "richresults verb"; a new verb + means a 2nd POST to the same endpoint for a payload already received. + Extended inspect() instead. +- **H1 moved up** (was AXE 5) — it is a PREREQUISITE of C1, not a + follow-up. Today only $DOMAIN (user-typed) is interpolated. After C1, N + URLs from a REMOTE sitemap flow into shell commands and fetch targets. + +### W2 (Bing) — DEFERRED, blocked on a real-world test +Killed after 4 challenge rounds. User's model: client sites live on CLIENT +Bing accounts, so a per-user API key means one key per client account. +OAuth is the right model but is a swamp: +- Redirect URI rejects ALL local forms (http/https/127.0.0.1 — user tested) +- Refresh tokens are **rotated + single-use**, self-described non-compliant + with OAuth 2.0 → store rewrite on every call, AND our parallel + seo/geo dispatch would race the rotation → invalid_grant, dead token +- Undocumented "anti-forgery token" failure on refresh, unanswered on Q&A +- MS's own advisor recommends falling back to the API key +- Doc contradicts itself on grant_type and the token endpoint; no library +REVIVAL CONDITION: a client already on Bing adds the user as a Read-Only +user → test in ~10 min whether the single API key sees DELEGATED sites +(undocumented, nobody knows). If yes → W2 is cheap and clean (one key, +client-owned verification, revocable, read-only, zero OAuth). If no → dead. +Value forgone meanwhile: Bing/DDG/Ecosia query stats + index status + +first-party backlinks. Real but modest; C1 dwarfs it. + +## 2026-07-16 — PLAN seo/geo parity vs claude-seo (superseded by the STATUS above) Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0, 5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install (install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md` diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 38a0cab..c0d666b 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -244,8 +244,14 @@ the PERMISSIVE template from `ai-crawlers-2026.md`. ### Live verification `[FULL only]` +**Guard the domain before it reaches a shell — mandatory, not optional.** +`$DOMAIN` is interpolated inside double quotes below, where `$` and backtick +still execute. Run the guard FIRST and use only its output; non-zero exit → +STOP this step and report the refusal, never sanitise-and-retry. + ```bash -DOMAIN="" +DOMAIN="$(bash ~/.claude/lib/url-guard.sh host "")" || { + echo "STEP 4 aborted: domain refused by url-guard"; exit 2; } # Verify robots.txt served curl -s "https://$DOMAIN/robots.txt" | head -50 @@ -443,13 +449,23 @@ grep -rhoE '"sameAs"[^]]*\]' \ --include="*.html" --include="*.astro" --include="*.tsx" --include="*.jsx" \ --include="*.vue" --include="*.svelte" --include="*.php" --include="*.json" \ . 2>/dev/null \ - | grep -oE 'https?://[^"]+' | sort -u | while read -r U; do + | grep -oE 'https?://[^"]+' | sort -u | while read -r RAW; do + # These URLs come from the audited repo's JSON-LD, not from the operator: + # guard each one before it reaches curl. A refused entry is REPORTED, not + # skipped silently — an unguardable sameAs is itself a finding. + U="$(bash ~/.claude/lib/url-guard.sh url "$RAW" 2>/dev/null)" || { + printf 'REFUSED %s\n' "$RAW"; continue; } printf '%s %s\n' \ "$(curl -sIL -o /dev/null -w '%{http_code}' --max-time 10 "$U" 2>/dev/null || echo 000)" \ "$U" done ``` +`REFUSED` rows are not dead links and not live ones — the URL never left the +machine. Report them in §14 with the raw value: a `sameAs` carrying shell +metacharacters or pointing at `localhost` is either broken markup or someone +probing, and both are worth the client knowing. + **Read the codes honestly — a block is not a death.** Some platforms refuse non-browser clients: LinkedIn answers `999` (verified 2026-07-16 against a live company page). A naive check calls that dead and the bundle deletes a diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 8c7070d..53a4bec 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -250,8 +250,15 @@ the §14 observed-list. But under `/seo` the security headers themselves are out of scope for scoring: see the Technical axis note in STEP 9. Under `/harden` they are the entire job. Reading is not scoring. +**Guard the domain before it reaches a shell — mandatory, not optional.** +Every curl below interpolates `$DOMAIN` inside double quotes, where `$` and +backtick still execute. Run the guard FIRST and use only its output; if it +exits non-zero, STOP this step and report the refusal — never "clean up" the +value and retry. + ```bash -DOMAIN="" +DOMAIN="$(bash ~/.claude/lib/url-guard.sh host "")" || { + echo "STEP 4 aborted: domain refused by url-guard"; exit 2; } # Headers curl -sI "https://$DOMAIN/" | head -30 diff --git a/lib/tests/url-guard.test.sh b/lib/tests/url-guard.test.sh new file mode 100644 index 0000000..679c832 --- /dev/null +++ b/lib/tests/url-guard.test.sh @@ -0,0 +1,74 @@ +#!/usr/bin/env bash +# lib/tests/url-guard.test.sh +set -u +G="$(cd "$(dirname "$0")/../.." && pwd)/lib/url-guard.sh" +pass=0; fail=0 +check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); + printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; } +# rc of a guard call, output discarded +rc() { bash "$G" "$1" "$2" >/dev/null 2>&1; return $?; } +# stdout of a guard call (empty on refusal) +out() { bash "$G" "$1" "$2" 2>/dev/null; } + +# --- hosts that must pass, echoing back unchanged --- +rc host "example.com"; check H1-plain "$?" 0 +rc host "www.sub.example.co.uk"; check H2-subdomains "$?" 0 +rc host "my-site.fr"; check H3-hyphen "$?" 0 +check H4-echoes-input "$(out host example.com)" "example.com" + +# --- shell metacharacters: the reason this guard exists --- +# Inside the double quotes seo-analyzer.md:257 uses, $ ` \ " break out. +rc host 'x$(id)'; check H5-cmdsubst "$?" 2 +rc host 'x`id`'; check H6-backtick "$?" 2 +rc host 'x;id'; check H7-semicolon "$?" 2 +rc host 'x|id'; check H8-pipe "$?" 2 +rc host 'x&id'; check H9-ampersand "$?" 2 +rc host 'x"'; check H10-dquote "$?" 2 +rc host "x'"; check H11-squote "$?" 2 +rc host 'x\y'; check H12-backslash "$?" 2 +rc host 'x y'; check H13-space "$?" 2 +rc host 'a +b'; check H14-newline "$?" 2 +# the real payload: read the OAuth vault into a request +rc host 'x$(cat ${HOME}/.claude/.env)'; check H15-env-exfil "$?" 2 +check H16-refusal-is-silent "$(out host 'x$(id)')" "" + +# --- literal local / private / metadata targets --- +rc host "localhost"; check L1-localhost "$?" 2 +rc host "LOCALHOST"; check L2-case-folded "$?" 2 +rc host "127.0.0.1"; check L3-loopback "$?" 2 +rc host "10.1.2.3"; check L4-private-10 "$?" 2 +rc host "192.168.1.1"; check L5-private-192 "$?" 2 +rc host "172.16.0.1"; check L6-private-172-lo "$?" 2 +rc host "172.31.255.254"; check L7-private-172-hi "$?" 2 +rc host "172.32.0.1"; check L8-172-32-is-public "$?" 0 +rc host "169.254.169.254"; check L9-link-local "$?" 2 +rc host "metadata.google.internal"; check L10-gcp-metadata "$?" 2 +rc host "0.0.0.0"; check L11-any-addr "$?" 2 +rc host "printer.local"; check L12-mdns "$?" 2 + +# --- urls --- +rc url "https://example.com/"; check U1-https "$?" 0 +rc url "http://example.com/a/b?x=1&y=2"; check U2-query "$?" 0 +rc url "https://example.com:8443/p"; check U3-port "$?" 0 +rc url "https://example.com/a%20b#frag"; check U4-pct-and-frag "$?" 0 +check U5-echoes-input "$(out url https://example.com/x)" "https://example.com/x" +rc url "ftp://example.com/"; check U6-ftp "$?" 2 +rc url "file:///etc/passwd"; check U7-file "$?" 2 +rc url "gopher://example.com/"; check U8-gopher "$?" 2 +rc url "example.com"; check U9-no-scheme "$?" 2 +rc url 'https://example.com/$(id)'; check U10-cmdsubst "$?" 2 +rc url 'https://example.com/`id`'; check U11-backtick "$?" 2 +rc url "https://localhost/x"; check U12-local "$?" 2 +rc url "https://127.0.0.1:8080/admin"; check U13-loopback "$?" 2 +# authority confusion: the real host is after the @, not before it +rc url "https://trusted.com@127.0.0.1/"; check U14-userinfo-local "$?" 2 +rc url "https://trusted.com@evil.com/"; check U15-userinfo-any "$?" 2 + +# --- usage --- +rc host ""; check X1-host-empty "$?" 2 +bash "$G" >/dev/null 2>&1; check X2-no-args "$?" 2 +bash "$G" bogus x >/dev/null 2>&1; check X3-bad-verb "$?" 2 +bash "$G" host a b >/dev/null 2>&1; check X4-extra-args "$?" 2 + +printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ] diff --git a/lib/url-guard.sh b/lib/url-guard.sh new file mode 100644 index 0000000..61a65cf --- /dev/null +++ b/lib/url-guard.sh @@ -0,0 +1,81 @@ +#!/usr/bin/env bash +# Validate a host or URL BEFORE it reaches a shell command or curl. +# Echoes the value on stdout when safe; exits 2 with a reason on stderr. +# +# HOST="$(bash ~/.claude/lib/url-guard.sh host "$RAW")" || exit 2 +# URL="$(bash ~/.claude/lib/url-guard.sh url "$RAW")" || exit 2 +# +# WHY: /seo and /geo interpolate externally-supplied strings into ~10 curl +# commands (seo-analyzer.md:254+, geo-analyzer.md:248+). Today $DOMAIN is typed +# by the operator, so the risk is self-inflicted. The sitemap crawl (C1) changes +# that: URLs then come from the TARGET'S OWN SERVER — a remote file whose bytes +# reach a shell. Inside the double quotes those curls use, the characters that +# break out are $ ` \ " — so a of +# https://x/$(cat ${HOME}/.claude/.env) +# would read GOOGLE_OAUTH_CLIENT_SECRET and CRUX_API_KEY straight out of the +# vault and into a request. Allowlist, per CLAUDE.md: explicit allowlist beats +# implicit denylist. +# +# NOT COVERED, deliberately: DNS-level SSRF. A public hostname that RESOLVES to +# a private address passes this guard. Closing that needs resolve-then-pin at +# the HTTP layer; curl in a shell cannot do it without a TOCTOU window between +# the check and the connection. Literal local targets ARE rejected below. The +# omission is stated rather than silent — see lib/seo-data/README.md. +set -uo pipefail + +_die() { echo "url-guard: $1" >&2; exit 2; } + +# Whole-string charset guards: C locale + POSIX `case`, the same shape as +# fetch.sh:25 _label_safe. Newline-proof and locale-independent, unlike a +# per-line grep. No `$` or backtick inside the patterns, so nothing expands. +_host_charset_ok() ( LC_ALL=C; case "$1" in + ''|[!A-Za-z0-9]*|*[!A-Za-z0-9.-]*) exit 1 ;; esac ) + +# Authority + path + query. Excludes $ ` \ " ' ; | ( ) * ! space and newline — +# none of which a real sitemap URL needs, all of which a shell reads. +_rest_charset_ok() ( LC_ALL=C; case "$1" in + ''|*[!A-Za-z0-9._~:/?#@=\&%+,-]*) exit 1 ;; esac ) + +# Literal local/private/metadata targets. This is a LITERAL check, not a DNS +# one: it stops the obvious, not a hostname that resolves inward. +_host_is_local() ( LC_ALL=C + # ${1,,} not tr: no fork, and no SC2018/SC2019 noise. Safe because the + # charset guard has already run — the string is [A-Za-z0-9.-] by here. + case "${1,,}" in + localhost|*.localhost|*.local|0.0.0.0|broadcasthost) exit 0 ;; + 127.*|10.*|169.254.*|192.168.*) exit 0 ;; + 172.1[6-9].*|172.2[0-9].*|172.3[01].*) exit 0 ;; + metadata.google.internal|metadata) exit 0 ;; + *) exit 1 ;; + esac ) + +_reject_local() { _host_is_local "$1" && _die "local/private target refused: '$1'"; return 0; } + +check_host() { + _host_charset_ok "$1" || _die "host charset (allowed A-Za-z0-9.-): '$1'" + _reject_local "$1" + printf '%s\n' "$1" +} + +check_url() { + local rest host + case "$1" in + https://*) rest="${1#https://}" ;; + http://*) rest="${1#http://}" ;; + *) _die "scheme must be http or https: '$1'" ;; + esac + _rest_charset_ok "$rest" || _die "url charset: '$1'" + host="${rest%%/*}"; host="${host%%\?*}"; host="${host%%#*}" + # user@host hides the real target: https://trusted.com@127.0.0.1/ hits .0.0.1 + case "$host" in *@*) _die "userinfo in authority (confusion vector): '$1'" ;; esac + host="${host%%:*}" # drop :port before validating the host + _host_charset_ok "$host" || _die "host charset: '$host'" + _reject_local "$host" + printf '%s\n' "$1" +} + +case "${1:-}" in + host) [ $# -eq 2 ] || _die "usage: url-guard.sh host "; check_host "$2" ;; + url) [ $# -eq 2 ] || _die "usage: url-guard.sh url "; check_url "$2" ;; + *) _die "usage: url-guard.sh {host|url} " ;; +esac From 8dcdc661ce64e2f6df8f442ce454696f6653f968 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 09:50:57 +0200 Subject: [PATCH 04/14] =?UTF-8?q?fix(seo,geo):=20C1a=20=E2=80=94=20find=20?= =?UTF-8?q?sees=20build=20output,=20grep=20does=20not;=20the=20two=20disag?= =?UTF-8?q?reed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Surfaced by dogfooding on a real Astro repo instead of reading the spec. Claude Code installs a shell function routing `grep` to ugrep with `--ignore-files`, so grep honours .gitignore and never descends into a gitignored dist/. `find` honours nothing. seo-analyzer uses both. Measured on zenquality (Astro, dist/ gitignored but built locally), seo-analyzer.md:497 returned 92 images, 45 of them under dist/ — every asset listed twice, source and generated copy, byte-identical. Two live consequences: - "top 20 by size" was ~10 real images dressed as 20. - Batch C (`cwebp -q 80 -o .webp`) could target dist/og-image.png; the .webp lands in dist/ and the `npm run build` the dispatcher runs to VERIFY the fix erases it. Fix lands, verification passes, nothing survives, report says applied. lib/source-scope.sh separates source from build output, framework-aware. `public/` is deliberately NOT excluded by default: it is Astro/Vite/Next SOURCE and holds favicon.ico, apple-touch-icon.png and robots.txt — the very files STEP 4 curls. It is build output only for Hugo/Gatsby, detected from config (legacy config.toml alone is ambiguous, so it needs archetypes/ too). Blanket-excluding it would blind the audit to its own resource checks. findargs emits one token per line and MUST be consumed via a quoted array. I shipped a flat-string version first and the dogfood caught it: the shell globs */dist/* against the CWD and passes the matches to find as search paths, which turned 90 hits into 135 and kept every dist/ file. Both the header and a functional test now pin that. Also: no bundle item may target build output, in either agent. That is the real safety net — even if some future find leaks a dist path, the fix cannot land there. Fix the source that generates the artifact; if the source cannot be found, that is a finding, not a reason to patch the artifact. SCOPE CORRECTION: the proposal claimed grep was auditing 86 generated files instead of 9 templates. That was FALSE — the ugrep shim already skips them. Killed my own premise before coding it; the real bug is narrower and lives in find only. Do NOT "fix" the grep lines to match: they are already correct, and adding these exclusions there would drop public/. Verified: 90 -> 45 images on the real repo, 0 dist survivors, public/ preserved (favicon.ico still visible); 34 new assertions PASS / 0 FAIL; full suite green; shellcheck clean. Test file addition used the documented one-shot config-edit sentinel. --- agents/geo-analyzer.md | 11 ++++- agents/seo-analyzer.md | 40 +++++++++++++++-- lib/source-scope.sh | 79 ++++++++++++++++++++++++++++++++++ lib/tests/source-scope.test.sh | 76 ++++++++++++++++++++++++++++++++ 4 files changed, 201 insertions(+), 5 deletions(-) create mode 100644 lib/source-scope.sh create mode 100644 lib/tests/source-scope.test.sh diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index c0d666b..7e51bf4 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -545,7 +545,8 @@ sample of a 300-page site says nothing about the other 294. ```bash # Extract H1/H2/H3 from main pages to assess heading style -for f in index.html $(find . -maxdepth 3 -name "*.astro" -o -name "*.tsx" -o -name "*.md" -o -name "*.html" | head -10); do +mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs) # C1a: skip build output +for f in index.html $(find . "${FEXCL[@]}" -maxdepth 3 \( -name "*.astro" -o -name "*.tsx" -o -name "*.md" -o -name "*.html" \) | head -10); do echo "=== $f ===" grep -oE '<(h1|h2|h3)[^>]*>[^<]+|^#{1,3} .+' "$f" 2>/dev/null | head -20 done @@ -970,6 +971,14 @@ PROCHAINE ETAPE : NEVER `Write` on shared templates. `Write` is reserved for files you solely own: robots.txt, llms.txt, llms-full.txt. Full-template refactor → escalate as user action in §11. +- **NEVER emit a bundle item targeting build output (C1a).** No path under + `dist/ build/ .next/ .nuxt/ .output/ _site/ .astro/ .svelte-kit/ out/` — + run `bash ~/.claude/lib/source-scope.sh list` for the authoritative set. + Those files are regenerated: the `npm run build` the dispatcher runs to + VERIFY your fix is what erases it. The fix lands, verification passes, + nothing survives, and the report claims it was applied. Fix the SOURCE + template that generates the file. If you cannot find the source, that is + a finding — say so, do not patch the artifact. - **Respect PERMISSIVE/RESTRICTIVE choice.** geo-analyzer defaults to PERMISSIVE (GEO's goal is AI visibility). Only switch if the client explicitly flags premium/regulated content. diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 53a4bec..9de8ffe 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -181,8 +181,9 @@ keep the two consistent. ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null # SEO files ls robots.txt sitemap.xml sitemap-index.xml sitemap-images.xml sitemap-videos.xml 2>/dev/null -# Legal pages -find . -maxdepth 3 \( -iname "*mention*" -o -iname "*legal*" -o -iname "*confidentialite*" -o -iname "*privacy*" -o -iname "*cgv*" -o -iname "*cgu*" \) 2>/dev/null | head -10 +# Legal pages — source only (C1a: find ignores .gitignore, grep does not) +mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs) +find . "${FEXCL[@]}" -maxdepth 3 \( -iname "*mention*" -o -iname "*legal*" -o -iname "*confidentialite*" -o -iname "*privacy*" -o -iname "*cgv*" -o -iname "*cgu*" \) 2>/dev/null | head -10 # Analytics / trackers grep -rl "gtag\|GTM-\|analytics\|matomo\|_paq\|plausible\|umami" --include="*.html" --include="*.js" --include="*.tsx" --include="*.astro" --include="*.php" . 2>/dev/null | head -10 # Cookie consent / CMP @@ -493,10 +494,32 @@ grep -rE ']*>' --include="*.html" --include="*.astro" --include="*.tsx" - # Images missing dimensions (CLS risk) grep -rE ']*>' --include="*.html" --include="*.astro" --include="*.tsx" --include="*.jsx" --include="*.php" . 2>/dev/null | grep -vE 'width=|height=' | head -30 -# Check image asset sizes -find . -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" \) ! -path "./node_modules/*" ! -path "./.git/*" -printf "%s %p\n" 2>/dev/null | sort -rn | head -20 +# Check image asset sizes — source only, never build output (C1a) +mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs) +find . "${FEXCL[@]}" -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" \) -printf "%s %p\n" 2>/dev/null | sort -rn | head -20 ``` +**Why the guard, and why `find` specifically (C1a).** `grep` and `find` +disagree about this repo and you use both. Claude Code routes `grep` through +ugrep with `--ignore-files`, so it honours `.gitignore` and never descends +into a gitignored `dist/`. `find` honours nothing. Measured on a real Astro +repo: this command returned **92 images, 45 of them under `dist/`** — every +asset twice, source and generated copy, byte-identical. So "top 20 by size" +was ~10 real images dressed as 20, and a batch-C item +(`cwebp -q 80 -o .webp`) could target `dist/og-image.png`, whose +`.webp` the dispatcher's own `npm run build` then erases. The fix lands, +verification passes, nothing survives. + +`FEXCL` MUST be consumed as a quoted array. `find . $FEXCL …` lets the shell +glob `*/dist/*` against the CWD and hand the matches to find as search paths +— that made the same run return 135 hits and kept every `dist/` file. + +Do NOT add these exclusions to the `grep` lines: the shim already covers +them, `public/` is deliberately kept (it is Astro/Vite/Next SOURCE and holds +`favicon.ico`, `apple-touch-icon.png`, `robots.txt` — the very files STEP 4 +curls), and it is build output only for Hugo/Gatsby, which the script +detects. + Flag images over 100 KB as compression candidates. WebP/AVIF preferred over JPEG/PNG. @@ -1195,6 +1218,15 @@ PROCHAINE ETAPE : `Write` on shared templates. `Write` is reserved for files you solely own: sitemap.xml, .htaccess, legal pages, new city/service pages. Full-template refactor → escalate as user action in §11. +- **NEVER emit a bundle item targeting build output (C1a).** No path under + `dist/ build/ .next/ .nuxt/ .output/ _site/ .astro/ .svelte-kit/ out/` — + `bash ~/.claude/lib/source-scope.sh list` is the authoritative set. Those + files are regenerated: the `npm run build` the dispatcher runs to VERIFY + your fix is what erases it. The fix lands, verification passes, nothing + survives, and the report claims it was applied. This bites batch C hardest + (`cwebp -q 80 -o .webp` on a `dist/` asset writes a `.webp` the + next build deletes). Fix the SOURCE that generates the artifact; if you + cannot find it, that is a finding — say so, do not patch the artifact. - **Landing page protection.** Zero visible change except meta tags, footer links, JSON-LD, image optimization. - **Preserve existing valid SEO.** Don't rewrite correct tags. diff --git a/lib/source-scope.sh b/lib/source-scope.sh new file mode 100644 index 0000000..ad63839 --- /dev/null +++ b/lib/source-scope.sh @@ -0,0 +1,79 @@ +#!/usr/bin/env bash +# Emit the directory exclusions that separate SOURCE from BUILD OUTPUT. +# +# EXCL="$(bash ~/.claude/lib/source-scope.sh grep)" +# grep -rl "gtag" $EXCL --include="*.html" . # note: $EXCL unquoted +# +# mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs) +# find . "${FEXCL[@]}" -iname '*.jpg' -printf '%s %p\n' # quoted array! +# +# findargs emits ONE TOKEN PER LINE and MUST be consumed through a quoted +# array. A flat string does not work: `find . $FEXCL ...` lets the shell glob +# `*/dist/*` against the CWD before find ever sees it, and the matches are then +# passed as search PATHS. Measured on zenquality: that turned 90 hits into 135 +# and kept every dist/ file. The array form passes each token literally. +# +# WHY: grep and find disagree about what is in the repo, and seo-analyzer uses +# both. +# +# grep → Claude Code installs a shell function routing grep to ugrep with +# `--ignore-files`, i.e. .gitignore-aware. A gitignored dist/ is +# invisible to it when recursing from `.`. Verified 2026-07-17. +# find → knows nothing about .gitignore. It sees everything. +# +# So on zenquality (Astro, dist/ gitignored, built locally) the spec's image +# audit at seo-analyzer.md:497 returns 92 images of which 45 live in dist/ — +# every asset listed twice, source and generated copy, identical bytes. Two +# real consequences: +# 1. "top 20 by size" is half generated duplicates: ~10 real images audited +# while 20 are claimed. +# 2. Batch C (`cwebp -q 80 -o .webp`) can target dist/og-image.png. +# The .webp lands in dist/ and the `npm run build` that /seo runs to VERIFY +# the fix erases it. The fix lands, verification passes, nothing survives. +# +# The grep side is already safe by accident — do NOT "fix" it to match find. +# `grep` mode below is defence in depth for the cases the shim misses: a repo +# that COMMITS its build output (no .gitignore entry to honour), or a directory +# that is not a git repo at all. +# +# `public/` is deliberately NOT in the always-list: it is SOURCE for +# Astro/Vite/Next and holds the very files this audit checks — favicon.ico, +# apple-touch-icon.png, robots.txt, OG images. It is build OUTPUT only for +# Hugo and Gatsby, detected below. Blanket-excluding it would blind the audit +# to its own resource checks. +# +# Exclusions are by NAME, not path, so a monorepo's frontend/dist is caught +# exactly like a root ./dist. +set -uo pipefail + +_die() { echo "source-scope: $1" >&2; exit 2; } + +# Build output + tool caches. Never source. +ALWAYS=(node_modules .git dist build .next .nuxt .output _site .astro + .svelte-kit .cache out coverage .vercel .netlify .turbo) + +# public/ is output for exactly these two generators. +_public_is_output() { + find . -maxdepth 3 \( -name "gatsby-config.js" -o -name "gatsby-config.ts" \ + -o -name "gatsby-config.mjs" -o -name "hugo.toml" -o -name "hugo.yaml" \ + -o -name "hugo.json" \) 2>/dev/null | read -r _ && return 0 + # Hugo's legacy config.toml is ambiguous on its own — pair it with archetypes/ + [ -d ./archetypes ] && [ -f ./config.toml ] && return 0 + return 1 +} + +_list() { + printf '%s\n' "${ALWAYS[@]}" + _public_is_output && printf 'public\n' + return 0 +} + +case "${1:-}" in + list) _list ;; + # Safe unquoted: --exclude-dir=NAME carries no glob character. + grep) _list | while read -r d; do printf -- '--exclude-dir=%s ' "$d"; done; echo ;; + # One token per line — consume with mapfile + a QUOTED array, never a flat + # string (see header: the shell would glob */dist/* against the CWD). + findargs) _list | while read -r d; do printf '!\n-path\n*/%s/*\n' "$d"; done ;; + *) _die "usage: source-scope.sh {list|grep|findargs}" ;; +esac diff --git a/lib/tests/source-scope.test.sh b/lib/tests/source-scope.test.sh new file mode 100644 index 0000000..b674b6c --- /dev/null +++ b/lib/tests/source-scope.test.sh @@ -0,0 +1,76 @@ +#!/usr/bin/env bash +# lib/tests/source-scope.test.sh +set -u +S="$(cd "$(dirname "$0")/../.." && pwd)/lib/source-scope.sh" +pass=0; fail=0 +check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); + printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; } +# does `list` (run inside dir $1) contain the name $2? +listed() { ( cd "$1" && bash "$S" list 2>/dev/null | grep -qxF "$2" ) \ + && echo yes || echo no; } + +TMP="$(mktemp -d)" + +# --- always-excluded build output + caches --- +mkdir -p "$TMP/plain" +for d in node_modules .git dist build .next .nuxt .output _site .astro \ + .svelte-kit .cache out coverage .vercel .netlify .turbo; do + check "A-$d-listed" "$(listed "$TMP/plain" "$d")" yes +done + +# --- public/ is SOURCE by default: Astro/Vite/Next keep favicon.ico, +# apple-touch-icon.png and robots.txt there, and the audit checks them --- +check B1-public-kept-by-default "$(listed "$TMP/plain" public)" no + +# --- public/ is OUTPUT for Gatsby and Hugo only --- +mkdir -p "$TMP/gatsby"; : > "$TMP/gatsby/gatsby-config.js" +check B2-gatsby-js "$(listed "$TMP/gatsby" public)" yes +mkdir -p "$TMP/gatsby2"; : > "$TMP/gatsby2/gatsby-config.ts" +check B3-gatsby-ts "$(listed "$TMP/gatsby2" public)" yes +mkdir -p "$TMP/hugo"; : > "$TMP/hugo/hugo.toml" +check B4-hugo-toml "$(listed "$TMP/hugo" public)" yes +mkdir -p "$TMP/hugo2"; : > "$TMP/hugo2/hugo.yaml" +check B5-hugo-yaml "$(listed "$TMP/hugo2" public)" yes +# legacy config.toml alone is ambiguous (many tools use it) — needs archetypes/ +mkdir -p "$TMP/amb"; : > "$TMP/amb/config.toml" +check B6-config-toml-alone-is-ambiguous "$(listed "$TMP/amb" public)" no +mkdir -p "$TMP/hugo3/archetypes"; : > "$TMP/hugo3/config.toml" +check B7-config-toml-plus-archetypes "$(listed "$TMP/hugo3" public)" yes + +# --- grep mode: flags, and no glob character (safe unquoted) --- +G="$(cd "$TMP/plain" && bash "$S" grep)" +case "$G" in *--exclude-dir=dist*) check C1-grep-has-dist ok ok ;; + *) check C1-grep-has-dist "missing" ok ;; esac +case "$G" in *"*"*) check C2-grep-has-no-glob "has-glob" ok ;; + *) check C2-grep-has-no-glob ok ok ;; esac + +# --- findargs: one token per line, 3 tokens per dir --- +N="$(cd "$TMP/plain" && bash "$S" findargs | wc -l)" +D="$(cd "$TMP/plain" && bash "$S" list | wc -l)" +check D1-findargs-3-tokens-per-dir "$N" "$((D * 3))" +check D2-findargs-first-token "$(cd "$TMP/plain" && bash "$S" findargs | head -1)" '!' + +# --- FUNCTIONAL: the array form actually excludes build output --- +# A flat unquoted string does NOT work here: the shell globs */dist/* against +# the CWD and passes the matches to find as search paths. Measured on a real +# repo, that turned 90 hits into 135 and kept every dist/ file. +W="$TMP/work"; mkdir -p "$W/src" "$W/dist" "$W/public" "$W/node_modules" +: > "$W/src/a.png"; : > "$W/dist/a.png"; : > "$W/public/favicon.ico" +: > "$W/node_modules/dep.png" +cd "$W" || exit 1 +mapfile -t FEXCL < <(bash "$S" findargs) +check E1-excludes-dist "$(find . "${FEXCL[@]}" -name 'a.png' | grep -c '/dist/')" 0 +check E2-keeps-src "$(find . "${FEXCL[@]}" -name 'a.png' | grep -c '/src/')" 1 +check E3-excludes-nodem "$(find . "${FEXCL[@]}" -name '*.png' | grep -c 'node_modules')" 0 +# public/ survives: the audit's own resource checks live there +check E4-keeps-public "$(find . "${FEXCL[@]}" -name 'favicon.ico' | wc -l)" 1 +cd / || exit 1 + +# --- usage --- +bash "$S" >/dev/null 2>&1; check X1-no-args "$?" 2 +bash "$S" bogus >/dev/null 2>&1; check X2-bad-verb "$?" 2 +# `find` was renamed to `findargs` when the flat-string form proved unsafe +bash "$S" find >/dev/null 2>&1; check X3-old-find-verb-gone "$?" 2 + +rm -rf "$TMP" +printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ] From 2de58faa38e74cec19a40c581d115954013c24d6 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Fri, 17 Jul 2026 11:28:00 +0200 Subject: [PATCH 05/14] =?UTF-8?q?feat(seo-data):=20C1b=20=E2=80=94=20sitem?= =?UTF-8?q?ap=20verb,=20the=20denominator=20COVERAGE=20never=20had?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit I5 made a COVERAGE line mandatory in STEP 9 and told the agent to "count the URLs in sitemap.xml" without giving it a command. STEP 4 only ever did `curl … | head -50` — a preview, not a count. This closes that. fetch.sh sitemap --url … → {count, urls[], index, dropped}. Stdlib only (urllib + xml.etree + gzip): no auth, no Google, no venv, so it runs wherever the mock/degrade paths run. Follows one level, dedupes, strips whitespace, handles .xml.gz. Every cap REPORTS what it cut (children_skipped, truncated) rather than truncating silently — same rule as COVERAGE itself. PLAN CORRECTION: the proposal said the verb would "validate each URL via the H1 guard". Wrong. urllib fetches these, so nothing here reaches a shell and there is no injection surface to guard. The guard belongs at the point of use, where seo-analyzer interpolates a URL into curl — which is the contract the sameAs check already established. A second copy of url-guard here would only drift from the first. The module carries a garbage filter, named as such. SECURITY: the security-guidance hook asked for defusedxml. Taken seriously, not obeyed — it would drag a venv into a module whose whole point is being stdlib-only. Split the threat instead: xml.etree does NOT expand external entities (XXE is not the vector), but it IS billion-laughs-vulnerable, and the 20 MB read ceiling bounds the input, not the expansion. A sitemap NEVER has a DTD — sitemaps.org is then — so any doctype/entity is refused BEFORE parsing, with its own reason (unsafe_xml_dtd, distinct from parse_failed: it is a finding, not a glitch). Refusing the construct beats depending on parser internals. Fixture is a real billion-laughs payload. Verified against the live target, not just fixtures: zenquality's sitemap returns count=86, dropped=0, matching `grep -c ''` on the raw XML exactly. Dead URL → {"status":"degraded","reason":"fetch_failed"}, exit 0. seo-data 95 -> 110 pass, 0 fail; full suite green; shellcheck + py_compile clean. Note: no config-edit sentinel was needed after all — config-protection guards lib/tests, not lib/seo-data. I posted one, found it uncommitted-and-unconsumed afterwards, and removed it rather than leave an open one-shot gate lying around. Worth knowing: seo-data.test.sh is 110 assertions and is NOT covered by that hook, while lib/tests/*.test.sh is. --- agents/seo-analyzer.md | 38 +++- lib/seo-data/README.md | 26 +++ lib/seo-data/fetch.sh | 5 +- lib/seo-data/fixtures-sitemap-dtd/sitemap.xml | 10 ++ .../fixtures-sitemap-index/sitemap.xml | 5 + .../fixtures-sitemap-index/sitemap_child.xml | 5 + lib/seo-data/fixtures/sitemap.xml | 15 ++ lib/seo-data/seo-data.test.sh | 30 ++++ lib/seo-data/sitemap.py | 163 ++++++++++++++++++ 9 files changed, 290 insertions(+), 7 deletions(-) create mode 100644 lib/seo-data/fixtures-sitemap-dtd/sitemap.xml create mode 100644 lib/seo-data/fixtures-sitemap-index/sitemap.xml create mode 100644 lib/seo-data/fixtures-sitemap-index/sitemap_child.xml create mode 100644 lib/seo-data/fixtures/sitemap.xml create mode 100644 lib/seo-data/sitemap.py diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 9de8ffe..80759ce 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -444,12 +444,38 @@ Fetch rendered HTML. Extract and analyze: ## STEP 5 — ON-PAGE AUDIT `[both]` **Record the denominator BEFORE sampling.** This step samples; the report -says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the -`head -50` in STEP 4 is a preview, not a count). That count is the coverage -denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap -→ denominator unknown: say so, never let silence imply full coverage. On a -500-page site a 12-page sample is 2.4% — the On-page score is an -extrapolation from it, and the reader cannot know that unless you print it. +says "audit". On a 500-page site a 12-page sample is 2.4% — the On-page score +is an extrapolation from it, and the reader cannot know unless you print it. + +```bash +bash ~/.claude/lib/seo-data/fetch.sh sitemap --url "https://$DOMAIN/sitemap.xml" +``` + +Returns `{count, urls[], index, dropped, ...}` — the coverage denominator and +your sampling frame. It follows a `` one level, dedupes, strips +whitespace, and handles `.xml.gz`. No auth, no venv, no Google. + +Read it honestly: +- `count` → the denominator for the STEP 9 COVERAGE line. +- `dropped > 0` → entries that were not usable URLs. Worth a §14 line: a + sitemap emitting junk is a tooling finding. +- `children_failed > 0` or `children_skipped` → the frame is incomplete. Say + so; do NOT present a partial denominator as the total. +- `status: degraded` → denominator UNKNOWN. Print that, never let silence + imply full coverage. `reason: unsafe_xml_dtd` is not a glitch — a sitemap + carrying a DTD is broken tooling or a billion-laughs aimed at the auditor. + Report it as a finding. + +**Guard every URL before it reaches curl.** These come from the target's own +server, not from the operator — the one place in this audit where a remote +file's bytes flow into a shell: + +```bash +U="$(bash ~/.claude/lib/url-guard.sh url "$RAW_FROM_SITEMAP")" || continue +``` + +The verb applies a garbage filter, not that guard; the guard belongs at the +point of use (same contract as the sameAs check in geo-analyzer). ### Meta tags per page (sample 5-15 key pages) diff --git a/lib/seo-data/README.md b/lib/seo-data/README.md index c21b097..faf1792 100644 --- a/lib/seo-data/README.md +++ b/lib/seo-data/README.md @@ -98,6 +98,32 @@ fetch.sh inspect --account client-a --property … --url https://ex.com/page • errors/warnings count issue INSTANCES; issues[] is deduped — the same issueMessage repeats across every affected item. +fetch.sh sitemap --url https://ex.com/sitemap.xml + → {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0, + "urls":["https://ex.com/", …]} + → {"status":"ok","index":true,"children_total":4,"children_read":4, + "children_failed":0,"count":312,…} # , one level deep + → {"status":"degraded","reason":"fetch_failed"|"parse_failed"|"no_urls" + |"unsafe_xml_dtd"} + + No auth, no Google, no venv: stdlib only (urllib + xml.etree + gzip). + Gives STEP 9's COVERAGE line the denominator it was told to print and never + had, and STEP 5 a real sampling frame. Dedupes, strips whitespace, handles + .xml.gz. Caps: 50 children of an index, 50k URLs, 20 MB read — each cut is + REPORTED (children_skipped / truncated), never silent. + + • NOT a security boundary. urllib fetches these, so nothing here reaches a + shell. The CONSUMER interpolates them into curl, so seo-analyzer runs + lib/url-guard.sh at the point of use — same contract as the sameAs check. + A second copy of the guard here would only drift. + • `unsafe_xml_dtd`: a sitemap NEVER has a DTD (sitemaps.org is then + ). Any doctype/entity is refused BEFORE parsing. xml.etree + does not expand external entities, but it IS billion-laughs-vulnerable — + 1 KB expands to gigabytes, and the 20 MB read ceiling bounds the input, + not the expansion. Refusing the construct beats depending on parser + internals AND keeps this stdlib-only; defusedxml would drag in a venv for + a document type that has no legitimate DTD. + fetch.sh forget --label client-a → {"status":"ok","removed":true|false} # false = label wasn't in the store diff --git a/lib/seo-data/fetch.sh b/lib/seo-data/fetch.sh index ede0ee8..8ca859d 100644 --- a/lib/seo-data/fetch.sh +++ b/lib/seo-data/fetch.sh @@ -29,6 +29,9 @@ case "$cmd" in accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;; crux|queries|inspect) exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;; + # No auth, no Google: stdlib-only, runs even without the venv. + sitemap) + exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;; forget) # forget --label