feat(lib): H1 — url-guard, shell-injection + local-target refusal before curl

Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is
typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+,
geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the
threat model completely: URLs then come from the TARGET'S OWN SERVER, so a
remote file's bytes reach a shell.

The severe hazard is injection, not SSRF. Those curls quote with ", inside
which $ and backtick still execute, and ~/.claude/.env holds
GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A <loc> of
`https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The
test suite asserts exactly that payload is refused.

Code, not prose: a markdown instruction does not stop an injection. Mirrors
the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C
locale, POSIX case: newline-proof, locale-independent, no grep pitfall.
Allowlist over denylist per CLAUDE.md.

Covers: shell metacharacters; scheme (http/https only — no file:, gopher:);
literal loopback/private/link-local/metadata/.local; userinfo authority
confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com).

NOT covered, stated in the header rather than left silent: DNS-level SSRF. A
public hostname resolving to a private address passes. Closing it needs
resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU
window. Proportionate to the threat model — this runs on a workstation
auditing the operator's own client sites.

Wired at all three entry points: both agents' STEP 4 domain assignment, and
the W3 sameAs loop (whose URLs come from the audited repo, not the operator).
Refused sameAs rows report as REFUSED rather than vanish — neither dead nor
live, and an unguardable sameAs is itself a finding.

Note: writing the test file tripped the config-protection hook (test suite is
a guarded quality-gate). Used the documented one-shot sentinel with a reason
rather than working around the gate; it was consumed as designed.

Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite
green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the
health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded
against the real zenquality.fr domain (accepted) and the real exfil payload
(refused, exit 2).
This commit is contained in:
Bastien Chanot
2026-07-17 09:25:34 +02:00
parent a6d423b940
commit 7d6aa09faf
5 changed files with 230 additions and 4 deletions
+49 -1
View File
@@ -1,6 +1,54 @@
# TODO # TODO
## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage) ## 2026-07-17 — STATUS seo/geo parity (branch bugfix/seo-geo-integrity, 10 commits, UNMERGED)
PHASE 1 — integrity: **DONE 7/7**. I3 8b0c98c · I1 57c67f2 · I2 4ea2fb8 ·
I5 64f175f · I4 e70e1d6 · I6 9da1dec · I8 acd452b. Plus 9cd7b51 (A1+A2, two
process anomalies surfaced by dogfooding /harden at zenquality.fr from the
wrong CWD).
PHASE 2 — free wins: W3 fe93b79 · W1 a6d423b · **W2 DEFERRED** (see below).
NEXT: H1 (SSRF/injection guard) → C1 (sitemap crawl). Human merge gate: all
10 commits await review; nothing merged to develop.
### Plan corrections made while executing (the plan was wrong 4×)
- **B3 KILLED** — GSC Links API does not exist. Verified against the API
reference: Search Console v1 exposes exactly Search Analytics, Sitemaps,
Sites, URL Inspection. A subagent hallucinated it; I doubted it in the
plan and the doubt was right. Common Crawl is the ONLY free backlink
source → the 70/100 cap is mandatory, not optional.
- **I1 was an over-correction** — "Off-page has ZERO data" was overstated
(relayed from a subagent, unverified). Brand mentions ARE gathered
(STEP 6). Narrowed the axis definition instead of N/A-ing it; weights
untouched to avoid churning historical scores twice.
- **I6 framing was wrong** — I claimed 3× that the stats "drive axis
weights". They do not; weight tables carry no citations. They drive Tier
recommendations and, worse, land in CLIENT reports via the "Cite sources"
rule. Reality was worse than my false version.
- **W1 was the wrong shape** — plan said "richresults verb"; a new verb
means a 2nd POST to the same endpoint for a payload already received.
Extended inspect() instead.
- **H1 moved up** (was AXE 5) — it is a PREREQUISITE of C1, not a
follow-up. Today only $DOMAIN (user-typed) is interpolated. After C1, N
URLs from a REMOTE sitemap flow into shell commands and fetch targets.
### W2 (Bing) — DEFERRED, blocked on a real-world test
Killed after 4 challenge rounds. User's model: client sites live on CLIENT
Bing accounts, so a per-user API key means one key per client account.
OAuth is the right model but is a swamp:
- Redirect URI rejects ALL local forms (http/https/127.0.0.1 — user tested)
- Refresh tokens are **rotated + single-use**, self-described non-compliant
with OAuth 2.0 → store rewrite on every call, AND our parallel
seo/geo dispatch would race the rotation → invalid_grant, dead token
- Undocumented "anti-forgery token" failure on refresh, unanswered on Q&A
- MS's own advisor recommends falling back to the API key
- Doc contradicts itself on grant_type and the token endpoint; no library
REVIVAL CONDITION: a client already on Bing adds the user as a Read-Only
user → test in ~10 min whether the single API key sees DELEGATED sites
(undocumented, nobody knows). If yes → W2 is cheap and clean (one key,
client-owned verification, revocable, read-only, zero OAuth). If no → dead.
Value forgone meanwhile: Bing/DDG/Ecosia query stats + index status +
first-party backlinks. Real but modest; C1 dwarfs it.
## 2026-07-16 — PLAN seo/geo parity vs claude-seo (superseded by the STATUS above)
Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0, Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0,
5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install 5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install
(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md` (install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md`
+18 -2
View File
@@ -244,8 +244,14 @@ the PERMISSIVE template from `ai-crawlers-2026.md`.
### Live verification `[FULL only]` ### Live verification `[FULL only]`
**Guard the domain before it reaches a shell — mandatory, not optional.**
`$DOMAIN` is interpolated inside double quotes below, where `$` and backtick
still execute. Run the guard FIRST and use only its output; non-zero exit →
STOP this step and report the refusal, never sanitise-and-retry.
```bash ```bash
DOMAIN="<production-domain>" DOMAIN="$(bash ~/.claude/lib/url-guard.sh host "<production-domain>")" || {
echo "STEP 4 aborted: domain refused by url-guard"; exit 2; }
# Verify robots.txt served # Verify robots.txt served
curl -s "https://$DOMAIN/robots.txt" | head -50 curl -s "https://$DOMAIN/robots.txt" | head -50
@@ -443,13 +449,23 @@ grep -rhoE '"sameAs"[^]]*\]' \
--include="*.html" --include="*.astro" --include="*.tsx" --include="*.jsx" \ --include="*.html" --include="*.astro" --include="*.tsx" --include="*.jsx" \
--include="*.vue" --include="*.svelte" --include="*.php" --include="*.json" \ --include="*.vue" --include="*.svelte" --include="*.php" --include="*.json" \
. 2>/dev/null \ . 2>/dev/null \
| grep -oE 'https?://[^"]+' | sort -u | while read -r U; do | grep -oE 'https?://[^"]+' | sort -u | while read -r RAW; do
# These URLs come from the audited repo's JSON-LD, not from the operator:
# guard each one before it reaches curl. A refused entry is REPORTED, not
# skipped silently — an unguardable sameAs is itself a finding.
U="$(bash ~/.claude/lib/url-guard.sh url "$RAW" 2>/dev/null)" || {
printf 'REFUSED %s\n' "$RAW"; continue; }
printf '%s %s\n' \ printf '%s %s\n' \
"$(curl -sIL -o /dev/null -w '%{http_code}' --max-time 10 "$U" 2>/dev/null || echo 000)" \ "$(curl -sIL -o /dev/null -w '%{http_code}' --max-time 10 "$U" 2>/dev/null || echo 000)" \
"$U" "$U"
done done
``` ```
`REFUSED` rows are not dead links and not live ones — the URL never left the
machine. Report them in §14 with the raw value: a `sameAs` carrying shell
metacharacters or pointing at `localhost` is either broken markup or someone
probing, and both are worth the client knowing.
**Read the codes honestly — a block is not a death.** Some platforms refuse **Read the codes honestly — a block is not a death.** Some platforms refuse
non-browser clients: LinkedIn answers `999` (verified 2026-07-16 against a non-browser clients: LinkedIn answers `999` (verified 2026-07-16 against a
live company page). A naive check calls that dead and the bundle deletes a live company page). A naive check calls that dead and the bundle deletes a
+8 -1
View File
@@ -250,8 +250,15 @@ the §14 observed-list. But under `/seo` the security headers themselves are
out of scope for scoring: see the Technical axis note in STEP 9. Under out of scope for scoring: see the Technical axis note in STEP 9. Under
`/harden` they are the entire job. Reading is not scoring. `/harden` they are the entire job. Reading is not scoring.
**Guard the domain before it reaches a shell — mandatory, not optional.**
Every curl below interpolates `$DOMAIN` inside double quotes, where `$` and
backtick still execute. Run the guard FIRST and use only its output; if it
exits non-zero, STOP this step and report the refusal — never "clean up" the
value and retry.
```bash ```bash
DOMAIN="<production-domain>" DOMAIN="$(bash ~/.claude/lib/url-guard.sh host "<production-domain>")" || {
echo "STEP 4 aborted: domain refused by url-guard"; exit 2; }
# Headers # Headers
curl -sI "https://$DOMAIN/" | head -30 curl -sI "https://$DOMAIN/" | head -30
+74
View File
@@ -0,0 +1,74 @@
#!/usr/bin/env bash
# lib/tests/url-guard.test.sh
set -u
G="$(cd "$(dirname "$0")/../.." && pwd)/lib/url-guard.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
# rc of a guard call, output discarded
rc() { bash "$G" "$1" "$2" >/dev/null 2>&1; return $?; }
# stdout of a guard call (empty on refusal)
out() { bash "$G" "$1" "$2" 2>/dev/null; }
# --- hosts that must pass, echoing back unchanged ---
rc host "example.com"; check H1-plain "$?" 0
rc host "www.sub.example.co.uk"; check H2-subdomains "$?" 0
rc host "my-site.fr"; check H3-hyphen "$?" 0
check H4-echoes-input "$(out host example.com)" "example.com"
# --- shell metacharacters: the reason this guard exists ---
# Inside the double quotes seo-analyzer.md:257 uses, $ ` \ " break out.
rc host 'x$(id)'; check H5-cmdsubst "$?" 2
rc host 'x`id`'; check H6-backtick "$?" 2
rc host 'x;id'; check H7-semicolon "$?" 2
rc host 'x|id'; check H8-pipe "$?" 2
rc host 'x&id'; check H9-ampersand "$?" 2
rc host 'x"'; check H10-dquote "$?" 2
rc host "x'"; check H11-squote "$?" 2
rc host 'x\y'; check H12-backslash "$?" 2
rc host 'x y'; check H13-space "$?" 2
rc host 'a
b'; check H14-newline "$?" 2
# the real payload: read the OAuth vault into a request
rc host 'x$(cat ${HOME}/.claude/.env)'; check H15-env-exfil "$?" 2
check H16-refusal-is-silent "$(out host 'x$(id)')" ""
# --- literal local / private / metadata targets ---
rc host "localhost"; check L1-localhost "$?" 2
rc host "LOCALHOST"; check L2-case-folded "$?" 2
rc host "127.0.0.1"; check L3-loopback "$?" 2
rc host "10.1.2.3"; check L4-private-10 "$?" 2
rc host "192.168.1.1"; check L5-private-192 "$?" 2
rc host "172.16.0.1"; check L6-private-172-lo "$?" 2
rc host "172.31.255.254"; check L7-private-172-hi "$?" 2
rc host "172.32.0.1"; check L8-172-32-is-public "$?" 0
rc host "169.254.169.254"; check L9-link-local "$?" 2
rc host "metadata.google.internal"; check L10-gcp-metadata "$?" 2
rc host "0.0.0.0"; check L11-any-addr "$?" 2
rc host "printer.local"; check L12-mdns "$?" 2
# --- urls ---
rc url "https://example.com/"; check U1-https "$?" 0
rc url "http://example.com/a/b?x=1&y=2"; check U2-query "$?" 0
rc url "https://example.com:8443/p"; check U3-port "$?" 0
rc url "https://example.com/a%20b#frag"; check U4-pct-and-frag "$?" 0
check U5-echoes-input "$(out url https://example.com/x)" "https://example.com/x"
rc url "ftp://example.com/"; check U6-ftp "$?" 2
rc url "file:///etc/passwd"; check U7-file "$?" 2
rc url "gopher://example.com/"; check U8-gopher "$?" 2
rc url "example.com"; check U9-no-scheme "$?" 2
rc url 'https://example.com/$(id)'; check U10-cmdsubst "$?" 2
rc url 'https://example.com/`id`'; check U11-backtick "$?" 2
rc url "https://localhost/x"; check U12-local "$?" 2
rc url "https://127.0.0.1:8080/admin"; check U13-loopback "$?" 2
# authority confusion: the real host is after the @, not before it
rc url "https://trusted.com@127.0.0.1/"; check U14-userinfo-local "$?" 2
rc url "https://trusted.com@evil.com/"; check U15-userinfo-any "$?" 2
# --- usage ---
rc host ""; check X1-host-empty "$?" 2
bash "$G" >/dev/null 2>&1; check X2-no-args "$?" 2
bash "$G" bogus x >/dev/null 2>&1; check X3-bad-verb "$?" 2
bash "$G" host a b >/dev/null 2>&1; check X4-extra-args "$?" 2
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
+81
View File
@@ -0,0 +1,81 @@
#!/usr/bin/env bash
# Validate a host or URL BEFORE it reaches a shell command or curl.
# Echoes the value on stdout when safe; exits 2 with a reason on stderr.
#
# HOST="$(bash ~/.claude/lib/url-guard.sh host "$RAW")" || exit 2
# URL="$(bash ~/.claude/lib/url-guard.sh url "$RAW")" || exit 2
#
# WHY: /seo and /geo interpolate externally-supplied strings into ~10 curl
# commands (seo-analyzer.md:254+, geo-analyzer.md:248+). Today $DOMAIN is typed
# by the operator, so the risk is self-inflicted. The sitemap crawl (C1) changes
# that: URLs then come from the TARGET'S OWN SERVER — a remote file whose bytes
# reach a shell. Inside the double quotes those curls use, the characters that
# break out are $ ` \ " — so a <loc> of
# https://x/$(cat ${HOME}/.claude/.env)
# would read GOOGLE_OAUTH_CLIENT_SECRET and CRUX_API_KEY straight out of the
# vault and into a request. Allowlist, per CLAUDE.md: explicit allowlist beats
# implicit denylist.
#
# NOT COVERED, deliberately: DNS-level SSRF. A public hostname that RESOLVES to
# a private address passes this guard. Closing that needs resolve-then-pin at
# the HTTP layer; curl in a shell cannot do it without a TOCTOU window between
# the check and the connection. Literal local targets ARE rejected below. The
# omission is stated rather than silent — see lib/seo-data/README.md.
set -uo pipefail
_die() { echo "url-guard: $1" >&2; exit 2; }
# Whole-string charset guards: C locale + POSIX `case`, the same shape as
# fetch.sh:25 _label_safe. Newline-proof and locale-independent, unlike a
# per-line grep. No `$` or backtick inside the patterns, so nothing expands.
_host_charset_ok() ( LC_ALL=C; case "$1" in
''|[!A-Za-z0-9]*|*[!A-Za-z0-9.-]*) exit 1 ;; esac )
# Authority + path + query. Excludes $ ` \ " ' ; | ( ) * ! space and newline —
# none of which a real sitemap URL needs, all of which a shell reads.
_rest_charset_ok() ( LC_ALL=C; case "$1" in
''|*[!A-Za-z0-9._~:/?#@=\&%+,-]*) exit 1 ;; esac )
# Literal local/private/metadata targets. This is a LITERAL check, not a DNS
# one: it stops the obvious, not a hostname that resolves inward.
_host_is_local() ( LC_ALL=C
# ${1,,} not tr: no fork, and no SC2018/SC2019 noise. Safe because the
# charset guard has already run — the string is [A-Za-z0-9.-] by here.
case "${1,,}" in
localhost|*.localhost|*.local|0.0.0.0|broadcasthost) exit 0 ;;
127.*|10.*|169.254.*|192.168.*) exit 0 ;;
172.1[6-9].*|172.2[0-9].*|172.3[01].*) exit 0 ;;
metadata.google.internal|metadata) exit 0 ;;
*) exit 1 ;;
esac )
_reject_local() { _host_is_local "$1" && _die "local/private target refused: '$1'"; return 0; }
check_host() {
_host_charset_ok "$1" || _die "host charset (allowed A-Za-z0-9.-): '$1'"
_reject_local "$1"
printf '%s\n' "$1"
}
check_url() {
local rest host
case "$1" in
https://*) rest="${1#https://}" ;;
http://*) rest="${1#http://}" ;;
*) _die "scheme must be http or https: '$1'" ;;
esac
_rest_charset_ok "$rest" || _die "url charset: '$1'"
host="${rest%%/*}"; host="${host%%\?*}"; host="${host%%#*}"
# user@host hides the real target: https://trusted.com@127.0.0.1/ hits .0.0.1
case "$host" in *@*) _die "userinfo in authority (confusion vector): '$1'" ;; esac
host="${host%%:*}" # drop :port before validating the host
_host_charset_ok "$host" || _die "host charset: '$host'"
_reject_local "$host"
printf '%s\n' "$1"
}
case "${1:-}" in
host) [ $# -eq 2 ] || _die "usage: url-guard.sh host <hostname>"; check_host "$2" ;;
url) [ $# -eq 2 ] || _die "usage: url-guard.sh url <url>"; check_url "$2" ;;
*) _die "usage: url-guard.sh {host|url} <value>" ;;
esac