From 8b0c98c99a242b0b0741a86c3fd6a2958488d9f7 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:06:30 +0200 Subject: [PATCH 1/8] =?UTF-8?q?fix(geo):=20I3=20=E2=80=94=20port=20NAP=20d?= =?UTF-8?q?irection=20rule=20(LRN-032)=20into=20geo-analyzer=20spec?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit geo-analyzer owns JSON-LD NAP (ownership matrix, seo/SKILL.md:261) and can rewrite it via G2 — AUTO tier, no confirmation (geo-analyzer.md:660). The LRN-032 protection lived ONLY in the /seo dispatcher prompt (seo/SKILL.md:339-343), so standalone /geo reconciled NAP with no canonical and no anti-seed guard — the exact zenquality trap, writing into client structured data. Root cause: a safety invariant that depended on the caller. Fixed at the layer that owns the data. - Data integrity: NAP direction rule, caller-independent, binds G2/G6. Covers CREATE (LocalBusiness from scratch) not just rewrite — geo builds missing schemas, seo-analyzer's wording only covered rewrite. - STEP 6 checklist: pointer at the line that triggers the action. Absent canonical is already the safe default (no directional fix), so no NAP collection step is needed in /geo — that would duplicate seo/SKILL.md STEP 0 and risk drift. Verified: make test 25+5+5 GREEN / 0 RED (incl. G3 strict-YAML frontmatter). --- .claude/tasks/TODO.md | 107 +++++++++++++++++++++++++++++++++++++++++ agents/geo-analyzer.md | 19 +++++++- 2 files changed, 125 insertions(+), 1 deletion(-) diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index feee292..2f78ca7 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,112 @@ # TODO +## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage) +Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0, +5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install +(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md` +deletes our seo-analyzer.md 42K it never installed; extensions/*/install.sh:42 +wipes settings.json on parse error; skills/seo/SKILL.md:119 injects Skool +upsell footer into deliverables). Their code is real (render_page.py 428 l +Playwright, url_safety.py 622 l SSRF, 326 tests, 320 pass) — adapt to our +fetch.sh contract, do NOT copy wholesale (no fail-open, no tokenstore, no +JSON shape). + +Framing: their plus-values map onto OUR integrity gaps — report claims more +than it measured. Same bar we held their README to. +Seam: `lib/seo-data/fetch.sh` verbs (accounts|crux|queries|inspect|forget) ++ fail-open `{"status":"degraded"}` + fixtures + tests. Everything below lands +as NEW VERBS. No new architecture. + +### AXE 0 — Integrity (no new deps, hours) — the score currently lies +- [ ] I1 Off-page axis scores 10-15% of FULL with ZERO data source (no API, + no index) → today fabricated, and it feeds /client-handover. Immediate + fix: extend existing LOCAL `N/A — requires FULL audit` pattern to FULL, + redistribute weights. Data upgrade later (AXE 3). Honesty now, data after. +- [ ] I2 VSI (Visual Stability Index) listed in CWV thresholds but NO path + retrieves it — neither CrUX nor PSI expose it. Phantom signal → remove + or source. +- [ ] I3 **SAFETY** /geo standalone: geo/SKILL.md (125 l) has no STEP 0, no + confirmed-NAP collection — but geo-analyzer OWNS JSON-LD NAP. Standalone + /geo on a local business can write unverified NAP with zero LRN-032 + protection. Real bug, not cosmetic. +- [ ] I4 Security headers counted 3× (seo-analyzer STEP 4 scores them in + Technical axis; depth-matrix.md says drop unless indexability; /harden + re-audits /100 with 3 validators). Contradiction between dedup rule and + agent spec → pick one owner. +- [ ] I5 Report says "audit", measured 5-15 sampled pages. State coverage % + explicitly in §0 until AXE 2 lands. + +### AXE 1 — Free wins on auth we ALREADY have (fetch.sh verbs) +- [ ] W1 `richresults` verb — GSC URL Inspection already returns + `richResultsResult`; our OAuth already carries the scope. Programmatic + rich-results validation on real Google data. **BEATS claude-seo**: their + README:314 "dual validator (Rich Results Test + Markup Validator)" is + FALSE — grep of all .py = zero calls, they are hyperlinks a human clicks. + Today our JSON-LD validity is LLM-read only. +- [ ] W2 `bing` verb — Bing Webmaster API, free. Closes the Google/Bing + asymmetry (Google = full OAuth layer, Bing = manual checklist) while + /geo targets ChatGPT Search, which indexes via Bing. Strategic, not cosmetic. +- [ ] W3 `sameas` resolution check — trivial curl loop. entity-seo.md lists + "sameAs pointing to dead profiles" as a known error class and never + checks it. ~10 lines. + +### AXE 2 — Coverage (biggest lever: ~97% of a 500-page site unseen today) +- [ ] C1 `crawl` verb — sitemap-driven URL discovery (we ALREADY fetch + sitemap.xml) + deterministic sampling + coverage % reported. No Chromium, + no paid API. Turns "5-15 LLM-chosen pages" into measured coverage. + Tradeoff vs claude-seo's link-following 500-page crawl: cheaper, but + misses unlinked/unsitemapped pages — accept + disclose. +- [ ] C2 Dupe/cannibalization detection — becomes possible once N pages in + hand: compare titles/H1/canonicals across the set. Free, unblocked by C1. +- [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated + as checks with no command to compute them. C1 unblocks real computation. + +### AXE 3 — Off-page real (upgrades I1) +- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph + (data.commoncrawl.org/projects/hyperlinkgraph), free, no key. +- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap + health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine + exactly. +- [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth + already there" — I doubt it: Search Console API v3 has no links endpoint + (links report is UI-only AFAIK). Verify before planning on it. Do not + assert. + +### AXE 4 — SPA blindness (dep decision — needs arbitrage) +- [ ] R1 `render` verb — Playwright, GATED on SPA detection (STEP 2 already + detects framework + rendering mode). Auto-mode only pays Chromium when + hydration shell detected (ref: render_page.py:226 logic, adapt not copy). +- [ ] R2 ARBITRAGE: heavy dep (Chromium ~300MB) vs our bash+curl purity. + Cheaper honest alternative: on SPA, REFUSE to score on-page rather than + score it wrong (today: curl reads source, not hydrated DOM → every + meta/JSON-LD/heading/img grep is blind, compensated only by a §0 flag). + +### AXE 5 — Hardening + regression (lower priority) +- [ ] H1 SSRF guard on curl paths — both agents curl user-supplied domains. + Our own CLAUDE.md doctrine says "never trust user input". url_safety.py + (622 l, obfuscated-IPv4 decode, DNS pinning) is a solid reference. +- [ ] H2 `drift` baseline (SQLite) — SEO.md Historique keeps only date+score+ + key changes. Their seo-drift is on-page regression detection, NOT rank + tracking (common misread). Optional. + +### NOT DOING (explicit, with reason) +- Keyword volumes → Google Ads Tier 3 needs ACTIVE ad spend (~$150-300/mo); + without spend the API returns buckets ("1K-10K"). Their own detect_tier() + never even returns 3 (google_auth.py:642-724 caps at 2) + google-ads absent + from requirements.txt. Not worth it. +- Real AI SoV (ChatGPT/Perplexity citation tracking) → paid everywhere + (SE Ranking/Profound/DataForSEO). Our current honest "not testable, here's + what we measured instead" disclosure BEATS faking it. Keep. +- Installing the plugin / +33 skills namespace → see destructive paths above. + +### Keep (already beats claude-seo — do not regress) +FR legal (LCEN/RGPD-ePrivacy/DGCCRF L121-1 — their whole repo: 2 hits, and +dma-consent-mode-v2.md:27 tells the agent to stay out) · fix-bundle + +ownership matrix + serial apply (their 18 agents are report-only, no +ownership discipline) · trajectory-to-17/20 + honest code ceiling (theirs is +flat 0-100, no legal axis) · llms.txt honest framing · NAP anti-dup-seed +(LRN-032). + ## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068) - [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop - [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index c9ce93c..6e14ead 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -360,7 +360,9 @@ action (G5 batch, confirmation needed — visible page creation). **Local business:** - [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.) -- [ ] NAP consistent with GMB +- [ ] NAP consistent with GMB — **direction rule applies** (Data integrity: + never pick a value from source majority; no canonical → no directional + fix) - [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable - [ ] `areaServed` lists served cities/regions - [ ] `openingHoursSpecification` matches reality @@ -895,6 +897,21 @@ PROCHAINE ETAPE : - **No invented entity data.** Never write a fake Wikidata QID, fake `sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown → placeholder `[À COMPLÉTER]` or omit. +- **NAP direction rule (LRN-032).** You own JSON-LD NAP, so this binds you + whoever called you — `/seo` passes a canonical, standalone `/geo` does + not. NEVER infer a correct NAP value from source majority: on-site + sources (JSON-LD, footer, settings DB, legal pages) usually descend from + ONE seed and can all carry the same wrong value — the single diverging + source may be the only one a human actually corrected. Direction of fix: + - Diverging from a CONFIRMED canonical field (passed by `/seo` STEP 0) + → fix the diverging source. + - Canonical UNCONFIRMED or absent (the standalone `/geo` case) → report + the divergence WITHOUT a directional fix; escalate as a user question + ("which value is correct?") in §11. + No G2/G6 item may write or rewrite a NAP value that no confirmed + canonical backs — **creating** a `LocalBusiness` from scratch included: + unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling + on-site source. - **Remove deprecated schemas rather than keep broken ones.** - **Cite sources.** When emitting stats in the report, link `content-shape-for-ai.md` research citations. From 57c67f2f7507c19ebc28b0d8131402279f789c90 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:14:12 +0200 Subject: [PATCH 2/8] =?UTF-8?q?fix(seo):=20I1=20=E2=80=94=20scope=20Off-pa?= =?UTF-8?q?ge=20axis=20to=20what=20is=20actually=20measured?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Axis was defined "backlinks, mentions, authority" (10% local / 15% national of the FULL score) but only mentions have a data source (STEP 6 web_search "" -site:). Backlinks and authority have no index, no API — the agent had to invent 2/3 of the number, and that number reaches a client via /client-handover. - Axis label names what is measured + points at §14. - Off-page axis note: score mentions ONLY; never price in unmeasured sub-components; a low mention count is NOT evidence of a weak backlink profile. Mandatory verbatim §14 line naming the gap + the nearest free source (Common Crawl) so the omission is legible, not silent. - LOCAL N/A label: was `N/A — requires FULL audit`, a promise FULL cannot keep for backlinks. Now states FULL covers brand mentions only. Weights deliberately unchanged: re-deriving now and again when a backlink source lands would churn historical scores twice. Revisit when the axis widens back (Common Crawl, phase 5). Note: initial plan was to mark the axis N/A in FULL and redistribute the weight. Reading the real spec (seo-analyzer.md:622 + STEP 6) showed that over-corrects — it discards the mentions data, which IS gathered. Narrowed the definition instead; composes with the Common Crawl work later. Verified: make test 35 GREEN / 0 RED. --- agents/seo-analyzer.md | 25 +++++++++++++++++++++++-- 1 file changed, 23 insertions(+), 2 deletions(-) diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index c60610a..43c36bb 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -619,7 +619,7 @@ FIX: AUTO () | USER () | Technical (perf, CWV, security headers, indexability) | 20% | 30% | | | On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | | | SEO Local (NAP, GMB, citations) | 25% | 5% | | -| Off-page (backlinks, mentions, authority) | 10% | 15% | | +| Off-page (unlinked brand mentions — backlinks/authority NOT auditable, §14) | 10% | 15% | | | Social presence | 10% | 5% | | | Competitive position | 5% | 10% | | | Legal compliance | 10% | 5% | | @@ -628,6 +628,23 @@ FIX: AUTO () | USER () real users, from STEP 4) when available; otherwise lab PageSpeed Lighthouse run. +**Off-page axis note (I1).** Score ONLY the unlinked brand mentions +gathered in STEP 6 (`web_search "" -site:`). +Backlink profile and domain authority have NO data source here — no index, +no API, nothing. NEVER price them into the number: an unmeasured +sub-component cannot be judged, and this axis carries 10-15% of a score +that reaches a client via `/client-handover`. A low mention count is a low +mention count — it is NOT evidence of a weak backlink profile. + +Mandatory §14 line whenever depth=FULL, verbatim: +`Backlinks / domain authority — NOT audited: no backlink index wired. +Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs / +Semrush / Majestic. The Off-page score above prices in brand mentions only.` + +Weight deliberately unchanged despite the narrower scope: re-deriving it +now, then again when a backlink source lands, would churn historical +scores twice. Revisit the 10/15% only when the axis widens back. + ### LOCAL depth — 4 axes | Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 | @@ -638,7 +655,11 @@ Lighthouse run. | Legal compliance (pages, CMP, mentions) | 20% | 15% | | LOCAL axes not audited (Off-page, Social, Competitive) appear as -`N/A — requires FULL audit` in the report. +`N/A — requires FULL audit` in the report. Off-page is the exception to +that promise: FULL audits its brand-mentions share ONLY — backlinks and +authority are unauditable at EVERY depth (see the Off-page axis note). +Print `N/A — FULL audits brand mentions only` for it, never a bare +"requires FULL audit" that FULL cannot keep. ### Projected code-only score + trajectory to 17/20 (mandatory) From 9cd7b51bb897f19ea7981fbbce7fec8ee3d0b6d1 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:25:16 +0200 Subject: [PATCH 3/8] =?UTF-8?q?fix(seo,geo):=20dogfood=20on=20zenquality.f?= =?UTF-8?q?r=20=E2=80=94=20two=20process=20anomalies?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Surfaced by pointing /harden at zenquality.fr from the claude-config CWD. A1 — no CWD/target coherence guard (systemic: /seo, /geo, /harden all lack it; grep confirms). A URL is supplied, the agent greps whatever CWD it landed in, nobody checks they are the same site. Demonstrated live: from claude-config, /harden would curl zenquality.fr while grepping claude-config, then score "Config hardening" on a codebase that is not the site. The live half looks right, the code half is fiction, and the report reads as authoritative. Fixed in both agents' STEP 2 rather than the 3 dispatchers: the agent does the grepping, so the guard binds whoever calls — same principle as I3. /harden inherits it free. A2 — seo-analyzer had zero origin-vs-edge awareness while geo-analyzer has the full CDN/WAF-override check (geo-analyzer.md:246-261). seo-analyzer does the infra detection AND is reused by /harden for its whole config-hardening axis (20/100). On zenquality — Apache origin behind a Scaleway nginx front — repo .htaccess + `server: nginx` invites the wrong call "nginx serves this, .htaccess is dead". I made that exact inference myself before reading the file. Rule added at STEP 2 infra detection: `server:` names the edge, not the origin; live-but-not-in-repo = "set upstream", never "missing". Verified: live probe of zenquality.fr (read-only, nothing written to the client repo); make test 35 GREEN / 0 RED. --- agents/geo-analyzer.md | 10 ++++++++++ agents/seo-analyzer.md | 28 ++++++++++++++++++++++++++++ 2 files changed, 38 insertions(+) diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 6e14ead..307e918 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -141,6 +141,16 @@ If called standalone via `/geo`, gather: ## STEP 2 — DETECT CONTEXT `[both]` +**FIRST — the CWD must BE the audited site.** You grep the current working +directory; no dispatcher checks that it matches the target domain. If a URL +was supplied and the CWD shows no web project at all (no `package.json` / +`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its +signals contradict the domain, STOP and report: +`CWD/TARGET MISMATCH — is not 's repo. Re-run from it, or +confirm live-only audit (LOCAL findings will be N/A).` +Never grep one codebase while curling another: the live half looks right, +the code half is fiction, and the report reads as authoritative. + ```bash # Framework (reuse detection from seo-analyzer if available) ls package.json composer.json Gemfile Cargo.toml go.mod 2>/dev/null diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 43c36bb..b5ef9bc 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -81,6 +81,17 @@ hreflang, infer from detected URL structures. ## STEP 2 — DETECT TECHNICAL CONTEXT `[both]` +**FIRST — the CWD must BE the audited site.** You grep the current working +directory; no dispatcher checks that it matches TARGET_URL. If a URL was +supplied and the CWD shows no web project at all (no `package.json` / +`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its +signals contradict the domain, STOP and report: +`CWD/TARGET MISMATCH — is not 's repo. Re-run from it, or +confirm live-only audit (LOCAL findings will be N/A).` +Never grep one codebase while curling another: the live half looks right, +the code half is fiction, and the report reads as authoritative. `/harden` +inherits this agent for its config axis, so the mismatch propagates there. + ### Framework & rendering ```bash @@ -148,6 +159,23 @@ RECOMMENDATION : KEEP & CONFIGURE plugin | INSTALL (P0 quick win) | M ### Infrastructure signals +**Origin vs edge — never infer the stack from `server:`.** That header names +whatever answered: usually the EDGE (Cloudflare, Scaleway/OVH front, CDN, +load balancer), not the origin. Apache behind an nginx front is a standard +topology — TLS terminated upstream, the origin sees plain HTTP plus +`X-Forwarded-Proto`. +- Repo `.htaccess` + `server: nginx` = NOT drift, NOT dead config. Do not + flag it, do not propose migrating it. +- Never move headers into an `nginx.conf` absent from the repo. Server-side + config you cannot read is a §14 gap, not a finding. +- A header present live but in no repo config = "set upstream", never + "missing". + +`/harden` reuses this agent for its entire config-hardening axis, so a wrong +topology call scores a client's server config against a file that never ran. +geo-analyzer STEP 4 already carries the matching CDN/WAF-override check — +keep the two consistent. + ```bash # Server / hosting ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null From 4ea2fb8c373aac0fc94a79a15933922872e23a55 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:46:52 +0200 Subject: [PATCH 4/8] =?UTF-8?q?fix(seo):=20I2=20=E2=80=94=20remove=20VSI,?= =?UTF-8?q?=20an=20SEO-blog=20fiction,=20from=20CWV=20thresholds?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit seo-analyzer.md:278 listed "VSI (Visual Stability Index) — new 2026 signal, Google Core Web Vitals 2.0" as a threshold, stated as fact, no hedge, in client-facing audits. It does not exist. Verified against two primary sources: - developer.chrome.com/docs/crux/api — complete metric list carries no visual_stability_index. Blogs claimed "Google is actively collecting VSI through CrUX": flatly false. - web.dev/articles/vitals — three stable CWV (LCP, INP, CLS). No VSI, no "Core Web Vitals 2.0". Thresholds change with prior notice on an annual cadence. Ten SEO blogs cross-cited each other into an apparent consensus. WebSearch returns that consensus, which is why the resources README rule "agents MUST cross-check via WebSearch on FULL" did not catch it — that mitigation launders blog misinformation into apparent verification. Fix removes the metric and states the sourcing rule where a future rumour would land: primary sources only (web.dev / Chromium blog / CrUX API list, the last being decisive — a metric CrUX cannot return is one we cannot score). Incident documented inline so it is not re-added. Verified: make test 35 GREEN / 0 RED. --- agents/seo-analyzer.md | 17 +++++++++++++++-- 1 file changed, 15 insertions(+), 2 deletions(-) diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index b5ef9bc..b00fdbd 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -275,8 +275,21 @@ Evaluate each present/missing: - **LCP** (Largest Contentful Paint) — < 2.5s - **INP** (Interaction to Next Paint) — < 200ms (replaced FID in Mar 2024) - **CLS** (Cumulative Layout Shift) — < 0.1 -- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web - Vitals 2.0 + +**Core Web Vitals are exactly these three** (web.dev/articles/vitals, +verified 2026-07-16). Google ships threshold changes with prior notice on a +predictable annual cadence — a "new CWV" that only SEO blogs know about does +not exist. Before adding a metric here, confirm it against a PRIMARY source: +web.dev, the Chromium blog, or `developer.chrome.com/docs/crux/api` — that +API metric list is decisive, because a metric CrUX cannot return is a metric +we cannot score. + +**WebSearch is not confirmation.** SEO blogs cross-cite each other into fake +consensus. A "VSI (Visual Stability Index) — new 2026 signal, Core Web +Vitals 2.0" line lived here until 2026-07-16 on exactly that basis: ten +blogs asserted it, several claimed CrUX was already collecting it, and it is +absent from both the CrUX API metric list and web.dev. Stated as fact, in a +threshold list, in client-facing audits. When a GSC account+property were passed in context, fetch CrUX field data first (**tilde path mandatory** — this agent runs from the From 64f175f01d50e388fef5e1d3bb516f3aa3be6747 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:50:19 +0200 Subject: [PATCH 5/8] =?UTF-8?q?fix(seo,geo):=20I5=20=E2=80=94=20disclose?= =?UTF-8?q?=20sampling=20coverage=20instead=20of=20implying=20an=20audit?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both agents sample (seo-analyzer.md:403 "sample 5-15 key pages", geo-analyzer.md:460 "Sample 5-10 key pages") and neither states it. The report says "audit". On a 500-page site a 12-page sample is 2.4%, and the reader cannot know that unless it is printed. /client-handover gates on these scores. The denominator was already within reach: STEP 4 fetches sitemap.xml. Count its URLs and the coverage ratio is free — same data C1 will use for sitemap-driven crawl later. - Mandatory COVERAGE line in both scoring blocks: N of M sitemap URLs (P%), or "total UNKNOWN" when no sitemap. Never omitted, never rounded up. - seo: <25% coverage repeats in §0 as a major alert. Sample by risk (one page per template + GSC position 4-10 quick wins), name skipped templates — an un-sampled template is an un-audited template. - geo: scoped honestly rather than blanket — COVERAGE bounds the per-page axes (Content Shape, page-level Schema.org) but NOT the site-wide ones (AI Crawlers Policy, llms.txt are single files, fully read). One ratio should not discredit axes it does not govern. Verified: make test 35 GREEN / 0 RED. --- agents/geo-analyzer.md | 13 +++++++++++++ agents/seo-analyzer.md | 22 ++++++++++++++++++++++ 2 files changed, 35 insertions(+) diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 307e918..c3e9e3b 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -459,6 +459,12 @@ Load: `~/.claude/agents/resources/content-shape-for-ai.md` Sample 5-10 key pages (homepage + top service/blog pages). For each: +**Record the denominator.** This samples; the report says "audit". Count the +URLs in `sitemap.xml` for the coverage ratio, and carry it into the GEO +SCORING block. No sitemap → total UNKNOWN, say so. Content shape is the +axis most damaged by silent sampling: it is judged per page, so a 6-page +sample of a 300-page site says nothing about the other 294. + ### Checks 1. **Definition Lead** — does the first sentence (or H1) follow @@ -586,6 +592,7 @@ Score each axis. Use concrete findings from STEP 2-9. ``` GEO SCORING () +COVERAGE : of sitemap URLs (

%) | pages, total UNKNOWN AI Crawlers Policy : XX/20 llms.txt : XX/20 Schema.org for AI : XX/20 @@ -596,6 +603,12 @@ AI Visibility (live) : XX/20 | N/A (LOCAL) GEO GLOBAL (weighted) : XX.X/20 () ``` +**COVERAGE is mandatory, never omitted, never rounded up.** It bounds the +per-page axes — Content Shape above all, and the page-level share of +Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected: +robots.txt and llms.txt are single files, fully read. Say which is which +rather than letting one ratio discredit the whole report. + Per user instruction: **GEO weight in combined SEO+GEO report = 20% for local, 25% for national/SaaS/content.** diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index b00fdbd..4536e6e 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -400,8 +400,22 @@ Fetch rendered HTML. Extract and analyze: ## STEP 5 — ON-PAGE AUDIT `[both]` +**Record the denominator BEFORE sampling.** This step samples; the report +says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the +`head -50` in STEP 4 is a preview, not a count). That count is the coverage +denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap +→ denominator unknown: say so, never let silence imply full coverage. On a +500-page site a 12-page sample is 2.4% — the On-page score is an +extrapolation from it, and the reader cannot know that unless you print it. + ### Meta tags per page (sample 5-15 key pages) +Sample by risk, not convenience: homepage + top templates (one per page +type: service, city, blog, product, legal) + any page GSC flags as a +position 4-10 quick win. Same template audited twice buys nothing; an +un-sampled template is an un-audited template — name the templates you +skipped. + For each sampled page: ``` PAGE: @@ -738,6 +752,8 @@ misroutes the client-handover gate and the user's effort. ``` SEO SCORING () +COVERAGE : of sitemap URLs (

%) — templates skipped: + | pages, total UNKNOWN (no sitemap) Technical : XX/20 On-page : XX/20 SEO Local : XX/20 | N/A @@ -749,6 +765,12 @@ Legal : XX/20 SEO GLOBAL (weighted): XX.X/20 () ``` +**COVERAGE is mandatory, never omitted, never rounded up.** It is the +honesty bound on every page-level axis: On-page and the on-page share of +Technical are extrapolations from the sample. If coverage < 25%, repeat it +in §0 as a major alert — a 17/20 drawn from 3% of a site is not a 17/20, and +`/client-handover` gates on these numbers. + Per user instruction: this score represents **80% of the combined final score for local B2C (20% for GEO), or 75% for SaaS/national (25% for GEO)**. The `/seo` dispatcher combines SEO and GEO scores. From e70e1d6c719838e3d80a81940ef08e128db72880 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 16:56:32 +0200 Subject: [PATCH 6/8] =?UTF-8?q?fix(seo):=20I4=20=E2=80=94=20stop=20double-?= =?UTF-8?q?counting=20security=20headers;=20/harden=20owns=20them?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Headers were scored three ways: seo-analyzer priced them into the Technical axis at both depths (:619 FULL, :635 LOCAL), depth-matrix.md:29 said drop them, and /harden re-audits them 0-100 against three external validators. The dedup rule and the agent spec contradicted each other; the agent won by default, so the same finding moved two scores in two reports. Arbitrated (user): /harden keeps them, /seo drops them. That confirms the rule that already existed — seo-analyzer was the violator. Constraint: /harden REUSES seo-analyzer, so the capability cannot be deleted, only scoped. Reading is not scoring: - Technical axis definitions no longer name security headers. - STEP 4 still curls them — needed for X-Robots-Tag, canonical/redirect coherence, and the §14 observed-list — but they earn no points under /seo. - Dispatched from /harden: unchanged, headers ARE the job (verified: its scope spec untouched, 16 header references intact). Carve-out: X-Robots-Tag stays in /seo under indexability. It is an indexing directive wearing a header's clothes — `noindex` there deindexes as surely as a meta robots tag. That is what depth-matrix.md:29 means by "unless it directly affects indexability"; the security headers do not. Drop is not silence: mandatory §14 line on FULL naming what was observed live plus a "run /harden " pointer. A user who never runs /harden must not read a clean Technical score as clean headers — same principle as the mandatory COVERAGE line (I5). Verified: make test 35 GREEN / 0 RED. --- agents/seo-analyzer.md | 35 +++++++++++++++++++++++++++++++++-- skills/seo/SKILL.md | 9 +++++++++ 2 files changed, 42 insertions(+), 2 deletions(-) diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 4536e6e..190a865 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -244,6 +244,12 @@ anonymous PageSpeed lab data and STEP 4/STEP 11 emit the §11 user action ### HTTP headers & security +**Read them; score them only for `/harden` (I4).** This section stays — the +raw headers are needed for `X-Robots-Tag`, canonical/redirect coherence, and +the §14 observed-list. But under `/seo` the security headers themselves are +out of scope for scoring: see the Technical axis note in STEP 9. Under +`/harden` they are the entire job. Reading is not scoring. + ```bash DOMAIN="" @@ -671,7 +677,7 @@ FIX: AUTO () | USER () | Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 | |---|---|---|---| -| Technical (perf, CWV, security headers, indexability) | 20% | 30% | | +| Technical (perf, CWV, indexability) | 20% | 30% | | | On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | | | SEO Local (NAP, GMB, citations) | 25% | 5% | | | Off-page (unlinked brand mentions — backlinks/authority NOT auditable, §14) | 10% | 15% | | @@ -683,6 +689,31 @@ FIX: AUTO () | USER () real users, from STEP 4) when available; otherwise lab PageSpeed Lighthouse run. +**Security headers are NOT scored here (I4).** `/harden` owns them and +grades them out of 100 with three external validators — pricing them into +this axis too was double-counting the same finding in two reports +(`depth-matrix.md:29` already said drop; this spec contradicted it). +- Dispatched from `/harden` (its prompt says NARROW-SCOPE): headers ARE the + job — audit and score them per its brief, ignore this note. +- Dispatched from `/seo`: do not score CSP, HSTS, X-Frame-Options, + X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP, + cookie flags. STEP 4 still reads them — you need them for the one + carve-out below — but they earn and lose no points here. + +**Carve-out — `X-Robots-Tag` stays.** It is an indexing directive wearing a +header's clothes: `noindex` served there deindexes the page as surely as a +meta robots tag. Score it under indexability. That is what +`depth-matrix.md:29` means by "unless it directly affects indexability" — +it is the header that does, and the security headers above are not. + +**Drop ≠ silence.** A user who never runs `/harden` must not read a clean +Technical score as clean headers. Whenever depth=FULL, emit in §14: +`Security headers (CSP, HSTS, X-Frame-Options…) — not scored here: /harden +owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden +. Observed live this run: .` +Name what you saw. An omission has to stay legible — the same reason +COVERAGE is mandatory in STEP 9. + **Off-page axis note (I1).** Score ONLY the unlinked brand mentions gathered in STEP 6 (`web_search "" -site:`). Backlink profile and domain authority have NO data source here — no index, @@ -704,7 +735,7 @@ scores twice. Revisit the 10/15% only when the axis widens back. | Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 | |---|---|---|---| -| Technical (security headers, indexability, config) | 25% | 35% | | +| Technical (indexability, config) | 25% | 35% | | | On-page (content, meta, headings, images, video, a11y, i18n) | 35% | 45% | | | SEO Local (markup, NAP in JSON-LD, legal) | 20% | 5% | | | Legal compliance (pages, CMP, mentions) | 20% | 15% | | diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index 6a5078d..99e9049 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -348,6 +348,15 @@ audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas, entity SEO, content shape for AI, AI visibility) — the geo-analyzer agent runs in parallel and owns those. +Do NOT score security headers either (CSP, HSTS, X-Frame-Options, +X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP, +cookie flags) — `/harden` owns them and grades them 0-100 against three +external validators (`depth-matrix.md:29`). Read them, keep +`X-Robots-Tag` under indexability (it is an indexing directive, not a +security header), and declare the rest in §14 with a "run /harden" pointer +plus what you observed live. Dropping them from the score must not make +them silent. + FILE OWNERSHIP (authoritative, prevents parallel-edit conflicts): - YOU OWN (read+write): sitemap.xml, image/video sitemaps, .htaccess, meta tags (title, description, OG, Twitter, canonical, robots meta), From 9da1dec9e6a360f71bf7de34690f7769f58eb520 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:32:55 +0200 Subject: [PATCH 7/8] =?UTF-8?q?fix(geo):=20I6=20=E2=80=94=20every=20stat?= =?UTF-8?q?=20was=20real=20and=20attached=20to=20the=20wrong=20claim?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Audited each statistic in agents/resources/ against primary sources after the VSI fiction (I2) showed WebSearch launders SEO-blog consensus. The failure mode is not invention — it is plausible recombination, which is what a model half-remembering a search result produces: - "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)" — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate over the whole method set, domain-dependent. No per-technique figure exists. - "Pages not updated quarterly are 3x more likely to lose AI citations (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source supports a quarterly decay multiplier. - "QAPage cited 58% more often than Article" — uncited. Nearest real number: AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources — wrong type, and its FAQPage figure (1.8%) points the opposite way to the claim it propped up. This one drove Tier 1 ranking. - "62% of searches involve voice" — uncited; 62% circulates as smart-speaker ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a 2014 Andrew Ng interview). Corrected my own framing too: I claimed three times these stats "drive axis weights". They do not — the weight tables carry no citations. They drive Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule pushed them into CLIENT reports as research-backed. Fixes: recommendations kept on mechanism, fabricated numbers removed with the incident documented inline so they are not re-added. Unverified stats (48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED] rather than asserted or deleted — I did not check them. Structural, not just exhortation: resources/README.md now mandates ` — — measured: — `. `measured:` is the field that catches this — all four errors survive a source name; none survives stating the real measurement next to the claim. WebSearch demoted from verification to crawler/tool-name lookup only. Verified: make test 35 GREEN / 0 RED. --- agents/geo-analyzer.md | 22 ++++++++--- agents/resources/README.md | 47 +++++++++++++++++++++++- agents/resources/ai-visibility-tools.md | 14 +++++-- agents/resources/content-shape-for-ai.md | 31 +++++++++++++--- agents/resources/geo-schemas.md | 28 ++++++++++++-- 5 files changed, 123 insertions(+), 19 deletions(-) diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index c3e9e3b..b3034db 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the ## Context — why GEO is its own discipline in 2026 -- AI Overviews trigger on ~48% of Google searches (April 2026). -- ChatGPT processes 2.5B queries/day. -- Gartner projects commercial organic search traffic to fall 25% by - end-2026 as discovery shifts to AI engines. +- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google + searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner + projects commercial organic search traffic to fall 25% by end-2026 as + discovery shifts to AI engines. Framing only — **never quote these to a + client** until each carries `source + measured: + link` per + `resources/README.md`. GEO is worth doing on mechanism; it does not need + these numbers to be true. - Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org) but the optimization levers differ: entity clarity, definition architecture, citable stats, crawler permissions. @@ -936,8 +939,15 @@ PROCHAINE ETAPE : unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling on-site source. - **Remove deprecated schemas rather than keep broken ones.** -- **Cite sources.** When emitting stats in the report, link - `content-shape-for-ai.md` research citations. +- **Cite sources, and only citable ones.** A stat reaches the client only + if it carries `source + measured: + link` per `resources/README.md`. + Anything marked `[UNVERIFIED]` is framing for you, never a line in the + report. Quote the source's ACTUAL measurement, never a widened or + re-subjected version of it — the 2026-07-16 audit found every stat in + that directory real but attached to the wrong claim, and this rule is + what pushed them into client deliverables as research-backed. + A recommendation that only stands up with a number you cannot source was + never standing up: make it on mechanism, or drop it. ### Process - **Every user action lists automation options.** Mandatory from diff --git a/agents/resources/README.md b/agents/resources/README.md index e685885..6283ea0 100644 --- a/agents/resources/README.md +++ b/agents/resources/README.md @@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current. These files capture state as of 2026-04. Crawler lists, Schema.org deprecations, and tool landscape shift fast. Agents MUST cross-check -via WebSearch on each run when FULL depth is selected. +crawler lists and tool names via WebSearch on each run when FULL depth is +selected. + +## Citation standard (mandatory for every statistic) + +**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO +blogs cross-cite each other into a consensus that looks like corroboration. +Two 2026-07-16 audits of this directory show how it fails: + +- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in + `seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already + collected it. It is absent from the CrUX API metric list and from + web.dev. WebSearch returned the echo, not the truth. +- Every stat in this directory was real **and attached to the wrong + subject**: the GEO paper's 40% (all methods) pinned on one technique; + LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay; + AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with + its meaning inverted; a smart-speaker adoption figure sold as voice-search + share. + +The failure mode is not invention — it is **plausible recombination**, which +is exactly what a model half-remembering a search result produces. So the +format has to make an unsourced number conspicuous: + +``` + — — measured: — +``` + +`measured:` is the field that catches it. All four errors above survive a +source name; none survives having to state the source's real measurement +next to the claim. + +Rules: +1. **Primary source or no number.** Peer-reviewed paper, the vendor's own + published study, or an official API/doc. `developer.chrome.com/docs/crux` + is decisive for metrics: what CrUX cannot return, we cannot score. +2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast, + Ahrefs publish useful data and sell products — say "vendor". +3. **Never widen scope.** An aggregate result is not a per-technique result. +4. **No number beats a wrong number.** A recommendation that only stands up + with a fabricated statistic was never standing up. Delete the stat, keep + the recommendation if it survives on mechanism. +5. **Unverified ⇒ labelled.** `[UNVERIFIED — ]` inline. Never quote an + unverified number to a client: `geo-analyzer.md` ("Cite sources") sends + these into client reports as research-backed. ## Loading pattern diff --git a/agents/resources/ai-visibility-tools.md b/agents/resources/ai-visibility-tools.md index 5f54efe..7c62e80 100644 --- a/agents/resources/ai-visibility-tools.md +++ b/agents/resources/ai-visibility-tools.md @@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI Overviews. -Context: Google AI Overviews trigger on ~48% of searches; ChatGPT -processes 2.5B queries/day; Gartner projects commercial organic -search traffic will drop 25% by 2026. Monitoring is no longer optional. +Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of +searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial +organic search traffic will drop 25% by 2026. + +> Not checked against primary sources in the 2026-07-16 audit that corrected +> the rest of this directory — flagged rather than asserted or deleted, per +> the citation standard in `README.md` (rule 5). The Gartner projection at +> least names its source; the other two float. Treat all three as +> motivation, not evidence: **do NOT quote them to a client** until each +> carries `source + measured: + link`. Their only job here is to explain why +> this file exists, and that argument does not need numbers. ## Commercial tools diff --git a/agents/resources/content-shape-for-ai.md b/agents/resources/content-shape-for-ai.md index 8c0f6a3..592eccc 100644 --- a/agents/resources/content-shape-for-ai.md +++ b/agents/resources/content-shape-for-ai.md @@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density. ### 4. Citations and statistics (strongest measured lever) -Adding peer-cited statistics with clear sources increases AI visibility -**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine -Optimization"). +Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024) +report that their optimisation methods **collectively** boost visibility +**by up to 40%** in generative-engine responses, and state the effect +**varies across domains**. Citations/statistics/quotations are among those +methods. + +> **Attribute this correctly.** Until 2026-07-16 this section read "Adding +> peer-cited statistics with clear sources increases AI visibility by up to +> 40%" — pinning the paper's *aggregate* result on this *one* technique. The +> paper publishes no separate figure per technique. When quoting it to a +> client: "up to 40%, across the method set, domain-dependent" — never "+40% +> if you add stats". Pattern: embed specific numbers with attribution. @@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure: ### 6. Freshness signals -Pages not updated at least quarterly are **3x more likely to lose AI -citations** (LLMRefs 2026 study). +Freshness is a real retrieval input: RAG systems fetch live and read +timestamps, so a page updated this quarter carries a stronger recency +signal than the same page last touched years ago. LLMrefs (a **vendor**, +not peer review) reports cited content running **~25.7% fresher** than +organic top-10 across ~17M citations. Substantive updates only — bumping a +date string is not freshness. + +> **The "3x" that lived here was grafted from another claim.** Until +> 2026-07-16 this read "Pages not updated at least quarterly are 3x more +> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x" +> says **brand mentions correlate ~3x more strongly with AI visibility than +> backlinks** — a different subject entirely. No source supports a quarterly +> decay multiplier. Recommend quarterly refresh on its merits; do not price +> it with a borrowed number. What to maintain: - Visible "Last updated: YYYY-MM-DD" at the top of content pages diff --git a/agents/resources/geo-schemas.md b/agents/resources/geo-schemas.md index da2f746..9d0eba9 100644 --- a/agents/resources/geo-schemas.md +++ b/agents/resources/geo-schemas.md @@ -21,8 +21,20 @@ existing instances. They no longer produce rich results. ### QAPage — single Q&A format -Pages cited 58% more often by ChatGPT vs basic Article schema. -Use when the page is built around ONE primary question. +Use when the page is built around ONE primary question. Emitting the type +that matches the content shape beats wrapping everything in a generic +`Article`. + +> **No lift figure here — the one that lived here was wrong.** Until +> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic +> Article schema", uncited. Nothing supports it. The nearest real number is +> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews / +> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%** +> of cited sources — a *prevalence* count for a *different type* — while +> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim +> it was propping up. Q&A shape is still worth doing on genuinely +> single-question pages; it is not worth a fabricated number. Do NOT quote a +> QAPage lift % to a client — there isn't one. ```json { @@ -81,8 +93,16 @@ visible content. ### Speakable — voice + AI extraction marker -62% of searches in 2026 involve voice. Speakable flags the passage -best suited for voice readout and AI summary. +Speakable flags the passage best suited for voice readout and AI summary. + +> **No voice-share figure — the one that lived here was a conflation.** +> Until 2026-07-16 this read "62% of searches in 2026 involve voice", +> uncited. No primary source carries it; 62% circulates as a *smart-speaker +> adoption* number, not a share of searches. It is the same family as the +> "50% of searches will be voice by 2020" myth — attributed to ComScore, +> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable +> is cheap and harmless, so keep recommending it on TL;DR / summary blocks — +> but justify it by extraction shape, never by a voice-share statistic. ```json { From acd452b92fa758bf3924f1aaeeb492c1c88c82c1 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:36:03 +0200 Subject: [PATCH 8/8] =?UTF-8?q?fix(seo):=20I8=20=E2=80=94=20drop=20the=20p?= =?UTF-8?q?hantom=20.claude/audits/external/=20precondition?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit STEP 0 told the user to run `mkdir -p .claude/audits/external` themselves before handing over an external report. Three things wrong with that: - The skill runs dozens of bash commands but outsourced this one to a human. - The timing was impossible: to "drop the export in" that directory the user needed it to already exist, so the instruction arrived after the moment it would have been useful. - The directory is not needed at all. `:218` already reads "File path given → Read it" — any path works — and nothing in skills/ or agents/ ever writes to that path. Grep confirms it is referenced by exactly these two lines and known to nothing else: a convention the skill invented, asked the user to create, and never used. Fix removes the precondition instead of automating it: give a path from anywhere, the tidy location stays a suggestion. Verified: make test 35 GREEN / 0 RED. --- skills/seo/SKILL.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index 99e9049..df561f3 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -204,9 +204,9 @@ Ask ONCE before dispatching the agents: ``` RAPPORT EXTERNE (optionnel) — un autre regard sur le site : - 1. Fichier — déposez l'export (PDF/MD/TXT) dans - `.claude/audits/external/` (ex. `sorank-YYYY-MM-DD.pdf`), - donnez le nom du fichier. (`mkdir -p .claude/audits/external`) + 1. Fichier — donnez le chemin de l'export (PDF/MD/TXT), où qu'il soit + (ex. `~/Téléchargements/sorank-2026-07-16.pdf`). Rangement conseillé + mais optionnel : `.claude/audits/external/`. 2. Collé — collez ici le contenu du PDF ou le "prompt pour IA" que l'outil suggère. 3. Ignorer — continuer sans. Le rapport final recommandera