diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index feee292..2f78ca7 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,112 @@ # TODO +## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage) +Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0, +5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install +(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md` +deletes our seo-analyzer.md 42K it never installed; extensions/*/install.sh:42 +wipes settings.json on parse error; skills/seo/SKILL.md:119 injects Skool +upsell footer into deliverables). Their code is real (render_page.py 428 l +Playwright, url_safety.py 622 l SSRF, 326 tests, 320 pass) — adapt to our +fetch.sh contract, do NOT copy wholesale (no fail-open, no tokenstore, no +JSON shape). + +Framing: their plus-values map onto OUR integrity gaps — report claims more +than it measured. Same bar we held their README to. +Seam: `lib/seo-data/fetch.sh` verbs (accounts|crux|queries|inspect|forget) ++ fail-open `{"status":"degraded"}` + fixtures + tests. Everything below lands +as NEW VERBS. No new architecture. + +### AXE 0 — Integrity (no new deps, hours) — the score currently lies +- [ ] I1 Off-page axis scores 10-15% of FULL with ZERO data source (no API, + no index) → today fabricated, and it feeds /client-handover. Immediate + fix: extend existing LOCAL `N/A — requires FULL audit` pattern to FULL, + redistribute weights. Data upgrade later (AXE 3). Honesty now, data after. +- [ ] I2 VSI (Visual Stability Index) listed in CWV thresholds but NO path + retrieves it — neither CrUX nor PSI expose it. Phantom signal → remove + or source. +- [ ] I3 **SAFETY** /geo standalone: geo/SKILL.md (125 l) has no STEP 0, no + confirmed-NAP collection — but geo-analyzer OWNS JSON-LD NAP. Standalone + /geo on a local business can write unverified NAP with zero LRN-032 + protection. Real bug, not cosmetic. +- [ ] I4 Security headers counted 3× (seo-analyzer STEP 4 scores them in + Technical axis; depth-matrix.md says drop unless indexability; /harden + re-audits /100 with 3 validators). Contradiction between dedup rule and + agent spec → pick one owner. +- [ ] I5 Report says "audit", measured 5-15 sampled pages. State coverage % + explicitly in §0 until AXE 2 lands. + +### AXE 1 — Free wins on auth we ALREADY have (fetch.sh verbs) +- [ ] W1 `richresults` verb — GSC URL Inspection already returns + `richResultsResult`; our OAuth already carries the scope. Programmatic + rich-results validation on real Google data. **BEATS claude-seo**: their + README:314 "dual validator (Rich Results Test + Markup Validator)" is + FALSE — grep of all .py = zero calls, they are hyperlinks a human clicks. + Today our JSON-LD validity is LLM-read only. +- [ ] W2 `bing` verb — Bing Webmaster API, free. Closes the Google/Bing + asymmetry (Google = full OAuth layer, Bing = manual checklist) while + /geo targets ChatGPT Search, which indexes via Bing. Strategic, not cosmetic. +- [ ] W3 `sameas` resolution check — trivial curl loop. entity-seo.md lists + "sameAs pointing to dead profiles" as a known error class and never + checks it. ~10 lines. + +### AXE 2 — Coverage (biggest lever: ~97% of a 500-page site unseen today) +- [ ] C1 `crawl` verb — sitemap-driven URL discovery (we ALREADY fetch + sitemap.xml) + deterministic sampling + coverage % reported. No Chromium, + no paid API. Turns "5-15 LLM-chosen pages" into measured coverage. + Tradeoff vs claude-seo's link-following 500-page crawl: cheaper, but + misses unlinked/unsitemapped pages — accept + disclose. +- [ ] C2 Dupe/cannibalization detection — becomes possible once N pages in + hand: compare titles/H1/canonicals across the set. Free, unblocked by C1. +- [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated + as checks with no command to compute them. C1 unblocks real computation. + +### AXE 3 — Off-page real (upgrades I1) +- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph + (data.commoncrawl.org/projects/hyperlinkgraph), free, no key. +- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap + health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine + exactly. +- [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth + already there" — I doubt it: Search Console API v3 has no links endpoint + (links report is UI-only AFAIK). Verify before planning on it. Do not + assert. + +### AXE 4 — SPA blindness (dep decision — needs arbitrage) +- [ ] R1 `render` verb — Playwright, GATED on SPA detection (STEP 2 already + detects framework + rendering mode). Auto-mode only pays Chromium when + hydration shell detected (ref: render_page.py:226 logic, adapt not copy). +- [ ] R2 ARBITRAGE: heavy dep (Chromium ~300MB) vs our bash+curl purity. + Cheaper honest alternative: on SPA, REFUSE to score on-page rather than + score it wrong (today: curl reads source, not hydrated DOM → every + meta/JSON-LD/heading/img grep is blind, compensated only by a §0 flag). + +### AXE 5 — Hardening + regression (lower priority) +- [ ] H1 SSRF guard on curl paths — both agents curl user-supplied domains. + Our own CLAUDE.md doctrine says "never trust user input". url_safety.py + (622 l, obfuscated-IPv4 decode, DNS pinning) is a solid reference. +- [ ] H2 `drift` baseline (SQLite) — SEO.md Historique keeps only date+score+ + key changes. Their seo-drift is on-page regression detection, NOT rank + tracking (common misread). Optional. + +### NOT DOING (explicit, with reason) +- Keyword volumes → Google Ads Tier 3 needs ACTIVE ad spend (~$150-300/mo); + without spend the API returns buckets ("1K-10K"). Their own detect_tier() + never even returns 3 (google_auth.py:642-724 caps at 2) + google-ads absent + from requirements.txt. Not worth it. +- Real AI SoV (ChatGPT/Perplexity citation tracking) → paid everywhere + (SE Ranking/Profound/DataForSEO). Our current honest "not testable, here's + what we measured instead" disclosure BEATS faking it. Keep. +- Installing the plugin / +33 skills namespace → see destructive paths above. + +### Keep (already beats claude-seo — do not regress) +FR legal (LCEN/RGPD-ePrivacy/DGCCRF L121-1 — their whole repo: 2 hits, and +dma-consent-mode-v2.md:27 tells the agent to stay out) · fix-bundle + +ownership matrix + serial apply (their 18 agents are report-only, no +ownership discipline) · trajectory-to-17/20 + honest code ceiling (theirs is +flat 0-100, no legal axis) · llms.txt honest framing · NAP anti-dup-seed +(LRN-032). + ## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068) - [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop - [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index c9ce93c..6e14ead 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -360,7 +360,9 @@ action (G5 batch, confirmation needed — visible page creation). **Local business:** - [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.) -- [ ] NAP consistent with GMB +- [ ] NAP consistent with GMB — **direction rule applies** (Data integrity: + never pick a value from source majority; no canonical → no directional + fix) - [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable - [ ] `areaServed` lists served cities/regions - [ ] `openingHoursSpecification` matches reality @@ -895,6 +897,21 @@ PROCHAINE ETAPE : - **No invented entity data.** Never write a fake Wikidata QID, fake `sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown → placeholder `[À COMPLÉTER]` or omit. +- **NAP direction rule (LRN-032).** You own JSON-LD NAP, so this binds you + whoever called you — `/seo` passes a canonical, standalone `/geo` does + not. NEVER infer a correct NAP value from source majority: on-site + sources (JSON-LD, footer, settings DB, legal pages) usually descend from + ONE seed and can all carry the same wrong value — the single diverging + source may be the only one a human actually corrected. Direction of fix: + - Diverging from a CONFIRMED canonical field (passed by `/seo` STEP 0) + → fix the diverging source. + - Canonical UNCONFIRMED or absent (the standalone `/geo` case) → report + the divergence WITHOUT a directional fix; escalate as a user question + ("which value is correct?") in §11. + No G2/G6 item may write or rewrite a NAP value that no confirmed + canonical backs — **creating** a `LocalBusiness` from scratch included: + unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling + on-site source. - **Remove deprecated schemas rather than keep broken ones.** - **Cite sources.** When emitting stats in the report, link `content-shape-for-ai.md` research citations.