forked from bchanot/claude
feat(seo-data): C1b — sitemap verb, the denominator COVERAGE never had
I5 made a COVERAGE line mandatory in STEP 9 and told the agent to "count the
URLs in sitemap.xml" without giving it a command. STEP 4 only ever did
`curl … | head -50` — a preview, not a count. This closes that.
fetch.sh sitemap --url … → {count, urls[], index, dropped}. Stdlib only
(urllib + xml.etree + gzip): no auth, no Google, no venv, so it runs wherever
the mock/degrade paths run. Follows <sitemapindex> one level, dedupes, strips
whitespace, handles .xml.gz. Every cap REPORTS what it cut (children_skipped,
truncated) rather than truncating silently — same rule as COVERAGE itself.
PLAN CORRECTION: the proposal said the verb would "validate each URL via the
H1 guard". Wrong. urllib fetches these, so nothing here reaches a shell and
there is no injection surface to guard. The guard belongs at the point of
use, where seo-analyzer interpolates a URL into curl — which is the contract
the sameAs check already established. A second copy of url-guard here would
only drift from the first. The module carries a garbage filter, named as such.
SECURITY: the security-guidance hook asked for defusedxml. Taken seriously,
not obeyed — it would drag a venv into a module whose whole point is being
stdlib-only. Split the threat instead: xml.etree does NOT expand external
entities (XXE is not the vector), but it IS billion-laughs-vulnerable, and
the 20 MB read ceiling bounds the input, not the expansion. A sitemap NEVER
has a DTD — sitemaps.org is <?xml?> then <urlset xmlns=> — so any
doctype/entity is refused BEFORE parsing, with its own reason
(unsafe_xml_dtd, distinct from parse_failed: it is a finding, not a glitch).
Refusing the construct beats depending on parser internals. Fixture is a real
billion-laughs payload.
Verified against the live target, not just fixtures: zenquality's sitemap
returns count=86, dropped=0, matching `grep -c '<loc>'` on the raw XML
exactly. Dead URL → {"status":"degraded","reason":"fetch_failed"}, exit 0.
seo-data 95 -> 110 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
Note: no config-edit sentinel was needed after all — config-protection guards
lib/tests, not lib/seo-data. I posted one, found it uncommitted-and-unconsumed
afterwards, and removed it rather than leave an open one-shot gate lying
around. Worth knowing: seo-data.test.sh is 110 assertions and is NOT covered
by that hook, while lib/tests/*.test.sh is.
This commit is contained in:
+32
-6
@@ -444,12 +444,38 @@ Fetch rendered HTML. Extract and analyze:
|
|||||||
## STEP 5 — ON-PAGE AUDIT `[both]`
|
## STEP 5 — ON-PAGE AUDIT `[both]`
|
||||||
|
|
||||||
**Record the denominator BEFORE sampling.** This step samples; the report
|
**Record the denominator BEFORE sampling.** This step samples; the report
|
||||||
says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the
|
says "audit". On a 500-page site a 12-page sample is 2.4% — the On-page score
|
||||||
`head -50` in STEP 4 is a preview, not a count). That count is the coverage
|
is an extrapolation from it, and the reader cannot know unless you print it.
|
||||||
denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap
|
|
||||||
→ denominator unknown: say so, never let silence imply full coverage. On a
|
```bash
|
||||||
500-page site a 12-page sample is 2.4% — the On-page score is an
|
bash ~/.claude/lib/seo-data/fetch.sh sitemap --url "https://$DOMAIN/sitemap.xml"
|
||||||
extrapolation from it, and the reader cannot know that unless you print it.
|
```
|
||||||
|
|
||||||
|
Returns `{count, urls[], index, dropped, ...}` — the coverage denominator and
|
||||||
|
your sampling frame. It follows a `<sitemapindex>` one level, dedupes, strips
|
||||||
|
whitespace, and handles `.xml.gz`. No auth, no venv, no Google.
|
||||||
|
|
||||||
|
Read it honestly:
|
||||||
|
- `count` → the denominator for the STEP 9 COVERAGE line.
|
||||||
|
- `dropped > 0` → entries that were not usable URLs. Worth a §14 line: a
|
||||||
|
sitemap emitting junk is a tooling finding.
|
||||||
|
- `children_failed > 0` or `children_skipped` → the frame is incomplete. Say
|
||||||
|
so; do NOT present a partial denominator as the total.
|
||||||
|
- `status: degraded` → denominator UNKNOWN. Print that, never let silence
|
||||||
|
imply full coverage. `reason: unsafe_xml_dtd` is not a glitch — a sitemap
|
||||||
|
carrying a DTD is broken tooling or a billion-laughs aimed at the auditor.
|
||||||
|
Report it as a finding.
|
||||||
|
|
||||||
|
**Guard every URL before it reaches curl.** These come from the target's own
|
||||||
|
server, not from the operator — the one place in this audit where a remote
|
||||||
|
file's bytes flow into a shell:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
U="$(bash ~/.claude/lib/url-guard.sh url "$RAW_FROM_SITEMAP")" || continue
|
||||||
|
```
|
||||||
|
|
||||||
|
The verb applies a garbage filter, not that guard; the guard belongs at the
|
||||||
|
point of use (same contract as the sameAs check in geo-analyzer).
|
||||||
|
|
||||||
### Meta tags per page (sample 5-15 key pages)
|
### Meta tags per page (sample 5-15 key pages)
|
||||||
|
|
||||||
|
|||||||
@@ -98,6 +98,32 @@ fetch.sh inspect --account client-a --property … --url https://ex.com/page
|
|||||||
• errors/warnings count issue INSTANCES; issues[] is deduped — the same
|
• errors/warnings count issue INSTANCES; issues[] is deduped — the same
|
||||||
issueMessage repeats across every affected item.
|
issueMessage repeats across every affected item.
|
||||||
|
|
||||||
|
fetch.sh sitemap --url https://ex.com/sitemap.xml
|
||||||
|
→ {"status":"ok","source":"sitemap","index":false,"count":86,"dropped":0,
|
||||||
|
"urls":["https://ex.com/", …]}
|
||||||
|
→ {"status":"ok","index":true,"children_total":4,"children_read":4,
|
||||||
|
"children_failed":0,"count":312,…} # <sitemapindex>, one level deep
|
||||||
|
→ {"status":"degraded","reason":"fetch_failed"|"parse_failed"|"no_urls"
|
||||||
|
|"unsafe_xml_dtd"}
|
||||||
|
|
||||||
|
No auth, no Google, no venv: stdlib only (urllib + xml.etree + gzip).
|
||||||
|
Gives STEP 9's COVERAGE line the denominator it was told to print and never
|
||||||
|
had, and STEP 5 a real sampling frame. Dedupes, strips whitespace, handles
|
||||||
|
.xml.gz. Caps: 50 children of an index, 50k URLs, 20 MB read — each cut is
|
||||||
|
REPORTED (children_skipped / truncated), never silent.
|
||||||
|
|
||||||
|
• NOT a security boundary. urllib fetches these, so nothing here reaches a
|
||||||
|
shell. The CONSUMER interpolates them into curl, so seo-analyzer runs
|
||||||
|
lib/url-guard.sh at the point of use — same contract as the sameAs check.
|
||||||
|
A second copy of the guard here would only drift.
|
||||||
|
• `unsafe_xml_dtd`: a sitemap NEVER has a DTD (sitemaps.org is <?xml?> then
|
||||||
|
<urlset xmlns=>). Any doctype/entity is refused BEFORE parsing. xml.etree
|
||||||
|
does not expand external entities, but it IS billion-laughs-vulnerable —
|
||||||
|
1 KB expands to gigabytes, and the 20 MB read ceiling bounds the input,
|
||||||
|
not the expansion. Refusing the construct beats depending on parser
|
||||||
|
internals AND keeps this stdlib-only; defusedxml would drag in a venv for
|
||||||
|
a document type that has no legitimate DTD.
|
||||||
|
|
||||||
fetch.sh forget --label client-a
|
fetch.sh forget --label client-a
|
||||||
→ {"status":"ok","removed":true|false} # false = label wasn't in the store
|
→ {"status":"ok","removed":true|false} # false = label wasn't in the store
|
||||||
|
|
||||||
|
|||||||
@@ -29,6 +29,9 @@ case "$cmd" in
|
|||||||
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
accounts) exec "$PY" "$HERE/tokenstore.py" list --file "$STORE" ;;
|
||||||
crux|queries|inspect)
|
crux|queries|inspect)
|
||||||
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
exec "$PY" "$HERE/google_seo.py" "$cmd" --store "$STORE" "$@" ;;
|
||||||
|
# No auth, no Google: stdlib-only, runs even without the venv.
|
||||||
|
sitemap)
|
||||||
|
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
|
||||||
forget)
|
forget)
|
||||||
# forget --label <label> → drop one account; forget --all → empty the store.
|
# forget --label <label> → drop one account; forget --all → empty the store.
|
||||||
# Local removal only — does NOT revoke the grant at Google's end.
|
# Local removal only — does NOT revoke the grant at Google's end.
|
||||||
@@ -41,6 +44,6 @@ case "$cmd" in
|
|||||||
fi
|
fi
|
||||||
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
||||||
exit 2 ;;
|
exit 2 ;;
|
||||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|forget} [flags]"}'
|
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|sitemap|forget} [flags]"}'
|
||||||
exit 2 ;;
|
exit 2 ;;
|
||||||
esac
|
esac
|
||||||
|
|||||||
@@ -0,0 +1,10 @@
|
|||||||
|
<?xml version="1.0"?>
|
||||||
|
<!DOCTYPE urlset [
|
||||||
|
<!ENTITY lol "lol">
|
||||||
|
<!ENTITY lol2 "&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;">
|
||||||
|
<!ENTITY lol3 "&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;">
|
||||||
|
<!ENTITY lol4 "&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;">
|
||||||
|
]>
|
||||||
|
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
|
||||||
|
<url><loc>https://ex.com/&lol4;</loc></url>
|
||||||
|
</urlset>
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
|
||||||
|
<sitemap><loc>https://ex.com/sitemap-pages.xml</loc></sitemap>
|
||||||
|
<sitemap><loc>https://ex.com/sitemap-blog.xml</loc></sitemap>
|
||||||
|
</sitemapindex>
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
|
||||||
|
<url><loc>https://ex.com/child-a</loc></url>
|
||||||
|
<url><loc>https://ex.com/child-b</loc></url>
|
||||||
|
</urlset>
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:xhtml="http://www.w3.org/1999/xhtml">
|
||||||
|
<url>
|
||||||
|
<loc>https://ex.com/</loc>
|
||||||
|
<changefreq>weekly</changefreq>
|
||||||
|
<xhtml:link rel="alternate" hreflang="en" href="https://ex.com/en/" />
|
||||||
|
</url>
|
||||||
|
<url><loc>https://ex.com/services</loc></url>
|
||||||
|
<url><loc>https://ex.com/blog</loc></url>
|
||||||
|
<url><loc>https://ex.com/blog</loc></url>
|
||||||
|
<url><loc> https://ex.com/spaced </loc></url>
|
||||||
|
<url><loc>ftp://ex.com/nope</loc></url>
|
||||||
|
<url><loc>https://ex.com/bad"quote</loc></url>
|
||||||
|
<url><loc></loc></url>
|
||||||
|
</urlset>
|
||||||
@@ -81,6 +81,36 @@ has "gsc degrades w/o creds" "$DEG" '"status": "degraded"'
|
|||||||
has "gsc degrade reason" "$DEG" 'no_credentials'
|
has "gsc degrade reason" "$DEG" 'no_credentials'
|
||||||
rm -rf "$TMP2"
|
rm -rf "$TMP2"
|
||||||
|
|
||||||
|
echo "── sitemap ──"
|
||||||
|
SM="$(SEO_DATA_MOCK_DIR="$MOCK" python3 "$SD/sitemap.py" --url https://ex.com/sitemap.xml)"
|
||||||
|
has "sitemap ok" "$SM" '"status": "ok"'
|
||||||
|
has "sitemap not an index" "$SM" '"index": false'
|
||||||
|
# fixture holds 8 <loc>: 1 empty, blog twice, ftp:// and a quoted one to drop
|
||||||
|
has "sitemap dedupes" "$SM" '"count": 4'
|
||||||
|
has "sitemap counts drops" "$SM" '"dropped": 2'
|
||||||
|
has "sitemap strips whitespace" "$SM" '"https://ex.com/spaced"'
|
||||||
|
hasnt "sitemap drops non-http" "$SM" 'ftp://'
|
||||||
|
hasnt "sitemap drops shell-meta" "$SM" 'bad"quote'
|
||||||
|
# namespace-agnostic: real sitemaps carry sitemaps.org xmlns (+ xhtml here)
|
||||||
|
has "sitemap reads namespaced" "$SM" '"https://ex.com/services"'
|
||||||
|
|
||||||
|
IDX="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-index" python3 "$SD/sitemap.py" \
|
||||||
|
--url https://ex.com/sitemap.xml)"
|
||||||
|
has "sitemapindex detected" "$IDX" '"index": true'
|
||||||
|
has "sitemapindex fans out" "$IDX" '"children_read": 2'
|
||||||
|
has "sitemapindex no child fail" "$IDX" '"children_failed": 0'
|
||||||
|
has "sitemapindex yields urls" "$IDX" '"https://ex.com/child-a"'
|
||||||
|
|
||||||
|
# A sitemap NEVER has a DTD. Refused at the door: xml.etree does not expand
|
||||||
|
# external entities but IS billion-laughs-vulnerable, and the 20MB read ceiling
|
||||||
|
# bounds the input, not the expansion. Refusing beats depending on the parser,
|
||||||
|
# and keeps this module stdlib-only (no defusedxml, no venv).
|
||||||
|
DTD="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-dtd" python3 "$SD/sitemap.py" \
|
||||||
|
--url https://ex.com/sitemap.xml)"
|
||||||
|
has "billion-laughs refused" "$DTD" '"status": "degraded"'
|
||||||
|
has "dtd reason is distinct" "$DTD" 'unsafe_xml_dtd'
|
||||||
|
hasnt "dtd never parsed" "$DTD" '"count"'
|
||||||
|
|
||||||
echo "── fetch.sh ──"
|
echo "── fetch.sh ──"
|
||||||
FETCH="$SD/fetch.sh"
|
FETCH="$SD/fetch.sh"
|
||||||
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
|
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
|
||||||
|
|||||||
@@ -0,0 +1,163 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Sitemap discovery -> normalized JSON. Stdlib only: no venv, no requests, no
|
||||||
|
auth. Gives STEP 9 COVERAGE the denominator it was told to report and never
|
||||||
|
had, and STEP 5 a real sampling frame instead of "5-15 key pages" chosen by
|
||||||
|
eye.
|
||||||
|
|
||||||
|
Deliberately NOT a security boundary. urllib fetches these URLs, so nothing
|
||||||
|
here reaches a shell and there is no injection surface to guard. The consumer
|
||||||
|
is different: seo-analyzer interpolates URLs into curl, so IT must run
|
||||||
|
lib/url-guard.sh at the point of use (same pattern as the sameAs check).
|
||||||
|
Duplicating the guard here would just add a second copy to drift. `_sane`
|
||||||
|
below is a cheap garbage filter, not that guard.
|
||||||
|
"""
|
||||||
|
import argparse, gzip, json, os
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
|
||||||
|
MAX_URLS = 50000 # sitemaps.org caps one file at 50k
|
||||||
|
MAX_CHILDREN = 50 # sitemapindex fan-out cap: bound the work, report the cut
|
||||||
|
TIMEOUT = 20
|
||||||
|
|
||||||
|
def _mock(name):
|
||||||
|
d = os.environ.get("SEO_DATA_MOCK_DIR")
|
||||||
|
if not d:
|
||||||
|
return None
|
||||||
|
path = os.path.join(d, name)
|
||||||
|
if not os.path.exists(path):
|
||||||
|
return None
|
||||||
|
with open(path, "rb") as f:
|
||||||
|
return f.read()
|
||||||
|
|
||||||
|
def _fetch(url):
|
||||||
|
from urllib.request import urlopen, Request # stdlib, lazy
|
||||||
|
req = Request(url, headers={"User-Agent": "claude-seo-data/1.0"})
|
||||||
|
with urlopen(req, timeout=TIMEOUT) as r: # nosec: audited target
|
||||||
|
raw = r.read(20 * 1024 * 1024) # 20 MB ceiling
|
||||||
|
if raw[:2] == b"\x1f\x8b": # sitemap.xml.gz is common
|
||||||
|
raw = gzip.decompress(raw)
|
||||||
|
return raw
|
||||||
|
|
||||||
|
class UnsafeXML(Exception):
|
||||||
|
"""A DTD reached the parser. Refused before parsing, not mitigated after."""
|
||||||
|
|
||||||
|
def _refuse_dtd(raw):
|
||||||
|
"""A sitemap NEVER has a DTD: sitemaps.org is <?xml?> then <urlset xmlns=>.
|
||||||
|
So refuse any doctype/entity outright, at the door.
|
||||||
|
|
||||||
|
This is the reason we do not pull in defusedxml. The stdlib parser is not
|
||||||
|
the problem for XXE — xml.etree.ElementTree does not expand external
|
||||||
|
entities, it raises on them — but it IS vulnerable to billion-laughs, where
|
||||||
|
a 1 KB document expands to gigabytes in RAM. The 20 MB read ceiling bounds
|
||||||
|
the input, not the expansion. Rejecting the construct beats depending on
|
||||||
|
the parser's internals, and keeps this module stdlib-only: no venv, same as
|
||||||
|
google_seo.py's mock/degrade paths. A sitemap with a DTD is not a sitemap
|
||||||
|
we want anyway.
|
||||||
|
"""
|
||||||
|
head = raw[:4096].lstrip()[:2048].upper()
|
||||||
|
if b"<!DOCTYPE" in head or b"<!ENTITY" in head:
|
||||||
|
raise UnsafeXML("DTD in sitemap")
|
||||||
|
|
||||||
|
def _locs(raw):
|
||||||
|
"""(<loc> texts, is_sitemapindex). Namespace-agnostic: real sitemaps carry
|
||||||
|
the sitemaps.org xmlns and often xhtml too."""
|
||||||
|
import xml.etree.ElementTree as ET # stdlib, lazy
|
||||||
|
_refuse_dtd(raw)
|
||||||
|
root = ET.fromstring(raw)
|
||||||
|
is_index = root.tag.endswith("sitemapindex")
|
||||||
|
out = []
|
||||||
|
for el in root.iter():
|
||||||
|
if el.tag.endswith("}loc") or el.tag == "loc":
|
||||||
|
text = (el.text or "").strip()
|
||||||
|
if text:
|
||||||
|
out.append(text)
|
||||||
|
return out, is_index
|
||||||
|
|
||||||
|
def _sane(u):
|
||||||
|
"""Cheap garbage filter — NOT lib/url-guard.sh. Drops what could never be a
|
||||||
|
real page URL; the consumer still guards before curling."""
|
||||||
|
if not u or len(u) > 2048:
|
||||||
|
return False
|
||||||
|
if any(c in u for c in '\n\r\t "\'\\`$<>{}|^'):
|
||||||
|
return False
|
||||||
|
return urlparse(u).scheme in ("http", "https")
|
||||||
|
|
||||||
|
def _expand(children):
|
||||||
|
"""Fetch each child sitemap of an index. A child that fails is skipped and
|
||||||
|
counted, never fatal: one dead child must not lose the other 49."""
|
||||||
|
urls, ok, failed = [], 0, 0
|
||||||
|
for c in children:
|
||||||
|
raw = _mock("sitemap_child.xml")
|
||||||
|
if raw is None:
|
||||||
|
try:
|
||||||
|
raw = _fetch(c)
|
||||||
|
except Exception:
|
||||||
|
failed += 1
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
sub, _ = _locs(raw)
|
||||||
|
except Exception:
|
||||||
|
failed += 1
|
||||||
|
continue
|
||||||
|
urls.extend(sub)
|
||||||
|
ok += 1
|
||||||
|
return urls, ok, failed
|
||||||
|
|
||||||
|
def sitemap(url):
|
||||||
|
raw = _mock("sitemap.xml")
|
||||||
|
if raw is None:
|
||||||
|
try:
|
||||||
|
raw = _fetch(url)
|
||||||
|
except Exception:
|
||||||
|
return {"status": "degraded", "reason": "fetch_failed"}
|
||||||
|
try:
|
||||||
|
locs, is_index = _locs(raw)
|
||||||
|
except UnsafeXML:
|
||||||
|
# Distinct from parse_failed on purpose: this one is a finding, not a
|
||||||
|
# glitch. A sitemap carrying a DTD is either broken tooling or someone
|
||||||
|
# aiming a billion-laughs at the auditor.
|
||||||
|
return {"status": "degraded", "reason": "unsafe_xml_dtd"}
|
||||||
|
except Exception:
|
||||||
|
return {"status": "degraded", "reason": "parse_failed"}
|
||||||
|
out = {"status": "ok", "source": "sitemap", "index": is_index}
|
||||||
|
if is_index:
|
||||||
|
out["children_total"] = len(locs)
|
||||||
|
kids, ok, failed = _expand(locs[:MAX_CHILDREN])
|
||||||
|
out["children_read"], out["children_failed"] = ok, failed
|
||||||
|
if len(locs) > MAX_CHILDREN: # say what was cut
|
||||||
|
out["children_skipped"] = len(locs) - MAX_CHILDREN
|
||||||
|
locs = kids
|
||||||
|
seen, urls, dropped = set(), [], 0
|
||||||
|
for u in locs:
|
||||||
|
if not _sane(u):
|
||||||
|
dropped += 1
|
||||||
|
continue
|
||||||
|
if u in seen:
|
||||||
|
continue
|
||||||
|
seen.add(u)
|
||||||
|
urls.append(u)
|
||||||
|
if len(urls) > MAX_URLS:
|
||||||
|
out["truncated"] = len(urls) - MAX_URLS
|
||||||
|
urls = urls[:MAX_URLS]
|
||||||
|
out["count"], out["dropped"], out["urls"] = len(urls), dropped, urls
|
||||||
|
if not urls:
|
||||||
|
return {"status": "degraded", "reason": "no_urls"}
|
||||||
|
return out
|
||||||
|
|
||||||
|
def _cli():
|
||||||
|
try:
|
||||||
|
p = argparse.ArgumentParser()
|
||||||
|
p.add_argument("--url", required=True)
|
||||||
|
p.add_argument("--store", default=None) # accepted+ignored: uniform dispatch
|
||||||
|
args = p.parse_args()
|
||||||
|
print(json.dumps(sitemap(args.url), indent=2))
|
||||||
|
except SystemExit as e:
|
||||||
|
if e.code not in (0, None):
|
||||||
|
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||||
|
raise
|
||||||
|
except Exception:
|
||||||
|
# Same fail-open contract as google_seo.py: never a traceback, never
|
||||||
|
# empty stdout, exit 0 so the audit degrades instead of dying.
|
||||||
|
print(json.dumps({"status": "degraded", "reason": "unexpected_error"}))
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
_cli()
|
||||||
Reference in New Issue
Block a user