Merge bugfix/seo-geo-integrity-phase1 into develop

This commit is contained in:
Bastien Chanot
2026-07-17 13:09:08 +02:00
8 changed files with 404 additions and 29 deletions
+107
View File
@@ -1,5 +1,112 @@
# TODO
## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage)
Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0,
5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install
(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md`
deletes our seo-analyzer.md 42K it never installed; extensions/*/install.sh:42
wipes settings.json on parse error; skills/seo/SKILL.md:119 injects Skool
upsell footer into deliverables). Their code is real (render_page.py 428 l
Playwright, url_safety.py 622 l SSRF, 326 tests, 320 pass) — adapt to our
fetch.sh contract, do NOT copy wholesale (no fail-open, no tokenstore, no
JSON shape).
Framing: their plus-values map onto OUR integrity gaps — report claims more
than it measured. Same bar we held their README to.
Seam: `lib/seo-data/fetch.sh` verbs (accounts|crux|queries|inspect|forget)
+ fail-open `{"status":"degraded"}` + fixtures + tests. Everything below lands
as NEW VERBS. No new architecture.
### AXE 0 — Integrity (no new deps, hours) — the score currently lies
- [ ] I1 Off-page axis scores 10-15% of FULL with ZERO data source (no API,
no index) → today fabricated, and it feeds /client-handover. Immediate
fix: extend existing LOCAL `N/A — requires FULL audit` pattern to FULL,
redistribute weights. Data upgrade later (AXE 3). Honesty now, data after.
- [ ] I2 VSI (Visual Stability Index) listed in CWV thresholds but NO path
retrieves it — neither CrUX nor PSI expose it. Phantom signal → remove
or source.
- [ ] I3 **SAFETY** /geo standalone: geo/SKILL.md (125 l) has no STEP 0, no
confirmed-NAP collection — but geo-analyzer OWNS JSON-LD NAP. Standalone
/geo on a local business can write unverified NAP with zero LRN-032
protection. Real bug, not cosmetic.
- [ ] I4 Security headers counted 3× (seo-analyzer STEP 4 scores them in
Technical axis; depth-matrix.md says drop unless indexability; /harden
re-audits /100 with 3 validators). Contradiction between dedup rule and
agent spec → pick one owner.
- [ ] I5 Report says "audit", measured 5-15 sampled pages. State coverage %
explicitly in §0 until AXE 2 lands.
### AXE 1 — Free wins on auth we ALREADY have (fetch.sh verbs)
- [ ] W1 `richresults` verb — GSC URL Inspection already returns
`richResultsResult`; our OAuth already carries the scope. Programmatic
rich-results validation on real Google data. **BEATS claude-seo**: their
README:314 "dual validator (Rich Results Test + Markup Validator)" is
FALSE — grep of all .py = zero calls, they are hyperlinks a human clicks.
Today our JSON-LD validity is LLM-read only.
- [ ] W2 `bing` verb — Bing Webmaster API, free. Closes the Google/Bing
asymmetry (Google = full OAuth layer, Bing = manual checklist) while
/geo targets ChatGPT Search, which indexes via Bing. Strategic, not cosmetic.
- [ ] W3 `sameas` resolution check — trivial curl loop. entity-seo.md lists
"sameAs pointing to dead profiles" as a known error class and never
checks it. ~10 lines.
### AXE 2 — Coverage (biggest lever: ~97% of a 500-page site unseen today)
- [ ] C1 `crawl` verb — sitemap-driven URL discovery (we ALREADY fetch
sitemap.xml) + deterministic sampling + coverage % reported. No Chromium,
no paid API. Turns "5-15 LLM-chosen pages" into measured coverage.
Tradeoff vs claude-seo's link-following 500-page crawl: cheaper, but
misses unlinked/unsitemapped pages — accept + disclose.
- [ ] C2 Dupe/cannibalization detection — becomes possible once N pages in
hand: compare titles/H1/canonicals across the set. Free, unblocked by C1.
- [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated
as checks with no command to compute them. C1 unblocks real computation.
### AXE 3 — Off-page real (upgrades I1)
- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph
(data.commoncrawl.org/projects/hyperlinkgraph), free, no key.
- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap
health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine
exactly.
- [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth
already there" — I doubt it: Search Console API v3 has no links endpoint
(links report is UI-only AFAIK). Verify before planning on it. Do not
assert.
### AXE 4 — SPA blindness (dep decision — needs arbitrage)
- [ ] R1 `render` verb — Playwright, GATED on SPA detection (STEP 2 already
detects framework + rendering mode). Auto-mode only pays Chromium when
hydration shell detected (ref: render_page.py:226 logic, adapt not copy).
- [ ] R2 ARBITRAGE: heavy dep (Chromium ~300MB) vs our bash+curl purity.
Cheaper honest alternative: on SPA, REFUSE to score on-page rather than
score it wrong (today: curl reads source, not hydrated DOM → every
meta/JSON-LD/heading/img grep is blind, compensated only by a §0 flag).
### AXE 5 — Hardening + regression (lower priority)
- [ ] H1 SSRF guard on curl paths — both agents curl user-supplied domains.
Our own CLAUDE.md doctrine says "never trust user input". url_safety.py
(622 l, obfuscated-IPv4 decode, DNS pinning) is a solid reference.
- [ ] H2 `drift` baseline (SQLite) — SEO.md Historique keeps only date+score+
key changes. Their seo-drift is on-page regression detection, NOT rank
tracking (common misread). Optional.
### NOT DOING (explicit, with reason)
- Keyword volumes → Google Ads Tier 3 needs ACTIVE ad spend (~$150-300/mo);
without spend the API returns buckets ("1K-10K"). Their own detect_tier()
never even returns 3 (google_auth.py:642-724 caps at 2) + google-ads absent
from requirements.txt. Not worth it.
- Real AI SoV (ChatGPT/Perplexity citation tracking) → paid everywhere
(SE Ranking/Profound/DataForSEO). Our current honest "not testable, here's
what we measured instead" disclosure BEATS faking it. Keep.
- Installing the plugin / +33 skills namespace → see destructive paths above.
### Keep (already beats claude-seo — do not regress)
FR legal (LCEN/RGPD-ePrivacy/DGCCRF L121-1 — their whole repo: 2 hits, and
dma-consent-mode-v2.md:27 tells the agent to stay out) · fix-bundle +
ownership matrix + serial apply (their 18 agents are report-only, no
ownership discipline) · trajectory-to-17/20 + honest code ceiling (theirs is
flat 0-100, no legal axis) · llms.txt honest framing · NAP anti-dup-seed
(LRN-032).
## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068)
- [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop
- [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail
+57 -7
View File
@@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the
## Context — why GEO is its own discipline in 2026
- AI Overviews trigger on ~48% of Google searches (April 2026).
- ChatGPT processes 2.5B queries/day.
- Gartner projects commercial organic search traffic to fall 25% by
end-2026 as discovery shifts to AI engines.
- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google
searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner
projects commercial organic search traffic to fall 25% by end-2026 as
discovery shifts to AI engines. Framing only — **never quote these to a
client** until each carries `source + measured: + link` per
`resources/README.md`. GEO is worth doing on mechanism; it does not need
these numbers to be true.
- Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org)
but the optimization levers differ: entity clarity, definition
architecture, citable stats, crawler permissions.
@@ -141,6 +144,16 @@ If called standalone via `/geo`, gather:
## STEP 2 — DETECT CONTEXT `[both]`
**FIRST — the CWD must BE the audited site.** You grep the current working
directory; no dispatcher checks that it matches the target domain. If a URL
was supplied and the CWD shows no web project at all (no `package.json` /
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
signals contradict the domain, STOP and report:
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
confirm live-only audit (LOCAL findings will be N/A).`
Never grep one codebase while curling another: the live half looks right,
the code half is fiction, and the report reads as authoritative.
```bash
# Framework (reuse detection from seo-analyzer if available)
ls package.json composer.json Gemfile Cargo.toml go.mod 2>/dev/null
@@ -360,7 +373,9 @@ action (G5 batch, confirmation needed — visible page creation).
**Local business:**
- [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.)
- [ ] NAP consistent with GMB
- [ ] NAP consistent with GMB — **direction rule applies** (Data integrity:
never pick a value from source majority; no canonical → no directional
fix)
- [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable
- [ ] `areaServed` lists served cities/regions
- [ ] `openingHoursSpecification` matches reality
@@ -447,6 +462,12 @@ Load: `~/.claude/agents/resources/content-shape-for-ai.md`
Sample 5-10 key pages (homepage + top service/blog pages). For each:
**Record the denominator.** This samples; the report says "audit". Count the
URLs in `sitemap.xml` for the coverage ratio, and carry it into the GEO
SCORING block. No sitemap → total UNKNOWN, say so. Content shape is the
axis most damaged by silent sampling: it is judged per page, so a 6-page
sample of a 300-page site says nothing about the other 294.
### Checks
1. **Definition Lead** — does the first sentence (or H1) follow
@@ -574,6 +595,7 @@ Score each axis. Use concrete findings from STEP 2-9.
```
GEO SCORING (<depth>)
COVERAGE : <N> of <M> sitemap URLs (<P>%) | <N> pages, total UNKNOWN
AI Crawlers Policy : XX/20 <justification>
llms.txt : XX/20 <justification>
Schema.org for AI : XX/20 <justification>
@@ -584,6 +606,12 @@ AI Visibility (live) : XX/20 | N/A (LOCAL)
GEO GLOBAL (weighted) : XX.X/20 (<depth>)
```
**COVERAGE is mandatory, never omitted, never rounded up.** It bounds the
per-page axes — Content Shape above all, and the page-level share of
Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected:
robots.txt and llms.txt are single files, fully read. Say which is which
rather than letting one ratio discredit the whole report.
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
local, 25% for national/SaaS/content.**
@@ -895,9 +923,31 @@ PROCHAINE ETAPE : <highest-priority>
- **No invented entity data.** Never write a fake Wikidata QID, fake
`sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown →
placeholder `[À COMPLÉTER]` or omit.
- **NAP direction rule (LRN-032).** You own JSON-LD NAP, so this binds you
whoever called you — `/seo` passes a canonical, standalone `/geo` does
not. NEVER infer a correct NAP value from source majority: on-site
sources (JSON-LD, footer, settings DB, legal pages) usually descend from
ONE seed and can all carry the same wrong value — the single diverging
source may be the only one a human actually corrected. Direction of fix:
- Diverging from a CONFIRMED canonical field (passed by `/seo` STEP 0)
→ fix the diverging source.
- Canonical UNCONFIRMED or absent (the standalone `/geo` case) → report
the divergence WITHOUT a directional fix; escalate as a user question
("which value is correct?") in §11.
No G2/G6 item may write or rewrite a NAP value that no confirmed
canonical backs — **creating** a `LocalBusiness` from scratch included:
unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling
on-site source.
- **Remove deprecated schemas rather than keep broken ones.**
- **Cite sources.** When emitting stats in the report, link
`content-shape-for-ai.md` research citations.
- **Cite sources, and only citable ones.** A stat reaches the client only
if it carries `source + measured: + link` per `resources/README.md`.
Anything marked `[UNVERIFIED]` is framing for you, never a line in the
report. Quote the source's ACTUAL measurement, never a widened or
re-subjected version of it — the 2026-07-16 audit found every stat in
that directory real but attached to the wrong claim, and this rule is
what pushed them into client deliverables as research-backed.
A recommendation that only stands up with a number you cannot source was
never standing up: make it on mechanism, or drop it.
### Process
- **Every user action lists automation options.** Mandatory from
+46 -1
View File
@@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current.
These files capture state as of 2026-04. Crawler lists, Schema.org
deprecations, and tool landscape shift fast. Agents MUST cross-check
via WebSearch on each run when FULL depth is selected.
crawler lists and tool names via WebSearch on each run when FULL depth is
selected.
## Citation standard (mandatory for every statistic)
**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO
blogs cross-cite each other into a consensus that looks like corroboration.
Two 2026-07-16 audits of this directory show how it fails:
- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in
`seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already
collected it. It is absent from the CrUX API metric list and from
web.dev. WebSearch returned the echo, not the truth.
- Every stat in this directory was real **and attached to the wrong
subject**: the GEO paper's 40% (all methods) pinned on one technique;
LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay;
AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with
its meaning inverted; a smart-speaker adoption figure sold as voice-search
share.
The failure mode is not invention — it is **plausible recombination**, which
is exactly what a model half-remembering a search result produces. So the
format has to make an unsourced number conspicuous:
```
<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>
```
`measured:` is the field that catches it. All four errors above survive a
source name; none survives having to state the source's real measurement
next to the claim.
Rules:
1. **Primary source or no number.** Peer-reviewed paper, the vendor's own
published study, or an official API/doc. `developer.chrome.com/docs/crux`
is decisive for metrics: what CrUX cannot return, we cannot score.
2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast,
Ahrefs publish useful data and sell products — say "vendor".
3. **Never widen scope.** An aggregate result is not a per-technique result.
4. **No number beats a wrong number.** A recommendation that only stands up
with a fabricated statistic was never standing up. Delete the stat, keep
the recommendation if it survives on mechanism.
5. **Unverified ⇒ labelled.** `[UNVERIFIED — <date>]` inline. Never quote an
unverified number to a client: `geo-analyzer.md` ("Cite sources") sends
these into client reports as research-backed.
## Loading pattern
+11 -3
View File
@@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
Overviews.
Context: Google AI Overviews trigger on ~48% of searches; ChatGPT
processes 2.5B queries/day; Gartner projects commercial organic
search traffic will drop 25% by 2026. Monitoring is no longer optional.
Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of
searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial
organic search traffic will drop 25% by 2026.
> Not checked against primary sources in the 2026-07-16 audit that corrected
> the rest of this directory — flagged rather than asserted or deleted, per
> the citation standard in `README.md` (rule 5). The Gartner projection at
> least names its source; the other two float. Treat all three as
> motivation, not evidence: **do NOT quote them to a client** until each
> carries `source + measured: + link`. Their only job here is to explain why
> this file exists, and that argument does not need numbers.
## Commercial tools
+26 -5
View File
@@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density.
### 4. Citations and statistics (strongest measured lever)
Adding peer-cited statistics with clear sources increases AI visibility
**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine
Optimization").
Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024)
report that their optimisation methods **collectively** boost visibility
**by up to 40%** in generative-engine responses, and state the effect
**varies across domains**. Citations/statistics/quotations are among those
methods.
> **Attribute this correctly.** Until 2026-07-16 this section read "Adding
> peer-cited statistics with clear sources increases AI visibility by up to
> 40%" — pinning the paper's *aggregate* result on this *one* technique. The
> paper publishes no separate figure per technique. When quoting it to a
> client: "up to 40%, across the method set, domain-dependent" — never "+40%
> if you add stats".
Pattern: embed specific numbers with attribution.
@@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure:
### 6. Freshness signals
Pages not updated at least quarterly are **3x more likely to lose AI
citations** (LLMRefs 2026 study).
Freshness is a real retrieval input: RAG systems fetch live and read
timestamps, so a page updated this quarter carries a stronger recency
signal than the same page last touched years ago. LLMrefs (a **vendor**,
not peer review) reports cited content running **~25.7% fresher** than
organic top-10 across ~17M citations. Substantive updates only — bumping a
date string is not freshness.
> **The "3x" that lived here was grafted from another claim.** Until
> 2026-07-16 this read "Pages not updated at least quarterly are 3x more
> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x"
> says **brand mentions correlate ~3x more strongly with AI visibility than
> backlinks** — a different subject entirely. No source supports a quarterly
> decay multiplier. Recommend quarterly refresh on its merits; do not price
> it with a borrowed number.
What to maintain:
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
+24 -4
View File
@@ -21,8 +21,20 @@ existing instances. They no longer produce rich results.
### QAPage — single Q&A format
Pages cited 58% more often by ChatGPT vs basic Article schema.
Use when the page is built around ONE primary question.
Use when the page is built around ONE primary question. Emitting the type
that matches the content shape beats wrapping everything in a generic
`Article`.
> **No lift figure here — the one that lived here was wrong.** Until
> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic
> Article schema", uncited. Nothing supports it. The nearest real number is
> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews /
> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%**
> of cited sources — a *prevalence* count for a *different type* — while
> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim
> it was propping up. Q&A shape is still worth doing on genuinely
> single-question pages; it is not worth a fabricated number. Do NOT quote a
> QAPage lift % to a client — there isn't one.
```json
{
@@ -81,8 +93,16 @@ visible content.
### Speakable — voice + AI extraction marker
62% of searches in 2026 involve voice. Speakable flags the passage
best suited for voice readout and AI summary.
Speakable flags the passage best suited for voice readout and AI summary.
> **No voice-share figure — the one that lived here was a conflation.**
> Until 2026-07-16 this read "62% of searches in 2026 involve voice",
> uncited. No primary source carries it; 62% circulates as a *smart-speaker
> adoption* number, not a share of searches. It is the same family as the
> "50% of searches will be voice by 2020" myth — attributed to ComScore,
> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable
> is cheap and harmless, so keep recommending it on TL;DR / summary blocks —
> but justify it by extraction shape, never by a voice-share statistic.
```json
{
+121 -6
View File
@@ -81,6 +81,17 @@ hreflang, infer from detected URL structures.
## STEP 2 — DETECT TECHNICAL CONTEXT `[both]`
**FIRST — the CWD must BE the audited site.** You grep the current working
directory; no dispatcher checks that it matches TARGET_URL. If a URL was
supplied and the CWD shows no web project at all (no `package.json` /
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
signals contradict the domain, STOP and report:
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
confirm live-only audit (LOCAL findings will be N/A).`
Never grep one codebase while curling another: the live half looks right,
the code half is fiction, and the report reads as authoritative. `/harden`
inherits this agent for its config axis, so the mismatch propagates there.
### Framework & rendering
```bash
@@ -148,6 +159,23 @@ RECOMMENDATION : KEEP & CONFIGURE plugin | INSTALL <plugin> (P0 quick win) | M
### Infrastructure signals
**Origin vs edge — never infer the stack from `server:`.** That header names
whatever answered: usually the EDGE (Cloudflare, Scaleway/OVH front, CDN,
load balancer), not the origin. Apache behind an nginx front is a standard
topology — TLS terminated upstream, the origin sees plain HTTP plus
`X-Forwarded-Proto`.
- Repo `.htaccess` + `server: nginx` = NOT drift, NOT dead config. Do not
flag it, do not propose migrating it.
- Never move headers into an `nginx.conf` absent from the repo. Server-side
config you cannot read is a §14 gap, not a finding.
- A header present live but in no repo config = "set upstream", never
"missing".
`/harden` reuses this agent for its entire config-hardening axis, so a wrong
topology call scores a client's server config against a file that never ran.
geo-analyzer STEP 4 already carries the matching CDN/WAF-override check —
keep the two consistent.
```bash
# Server / hosting
ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null
@@ -216,6 +244,12 @@ anonymous PageSpeed lab data and STEP 4/STEP 11 emit the §11 user action
### HTTP headers & security
**Read them; score them only for `/harden` (I4).** This section stays — the
raw headers are needed for `X-Robots-Tag`, canonical/redirect coherence, and
the §14 observed-list. But under `/seo` the security headers themselves are
out of scope for scoring: see the Technical axis note in STEP 9. Under
`/harden` they are the entire job. Reading is not scoring.
```bash
DOMAIN="<production-domain>"
@@ -247,8 +281,21 @@ Evaluate each present/missing:
- **LCP** (Largest Contentful Paint) — < 2.5s
- **INP** (Interaction to Next Paint) — < 200ms (replaced FID in Mar 2024)
- **CLS** (Cumulative Layout Shift) — < 0.1
- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web
Vitals 2.0
**Core Web Vitals are exactly these three** (web.dev/articles/vitals,
verified 2026-07-16). Google ships threshold changes with prior notice on a
predictable annual cadence — a "new CWV" that only SEO blogs know about does
not exist. Before adding a metric here, confirm it against a PRIMARY source:
web.dev, the Chromium blog, or `developer.chrome.com/docs/crux/api` — that
API metric list is decisive, because a metric CrUX cannot return is a metric
we cannot score.
**WebSearch is not confirmation.** SEO blogs cross-cite each other into fake
consensus. A "VSI (Visual Stability Index) — new 2026 signal, Core Web
Vitals 2.0" line lived here until 2026-07-16 on exactly that basis: ten
blogs asserted it, several claimed CrUX was already collecting it, and it is
absent from both the CrUX API metric list and web.dev. Stated as fact, in a
threshold list, in client-facing audits.
When a GSC account+property were passed in context, fetch CrUX field
data first (**tilde path mandatory** — this agent runs from the
@@ -359,8 +406,22 @@ Fetch rendered HTML. Extract and analyze:
## STEP 5 — ON-PAGE AUDIT `[both]`
**Record the denominator BEFORE sampling.** This step samples; the report
says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the
`head -50` in STEP 4 is a preview, not a count). That count is the coverage
denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap
→ denominator unknown: say so, never let silence imply full coverage. On a
500-page site a 12-page sample is 2.4% — the On-page score is an
extrapolation from it, and the reader cannot know that unless you print it.
### Meta tags per page (sample 5-15 key pages)
Sample by risk, not convenience: homepage + top templates (one per page
type: service, city, blog, product, legal) + any page GSC flags as a
position 4-10 quick win. Same template audited twice buys nothing; an
un-sampled template is an un-audited template — name the templates you
skipped.
For each sampled page:
```
PAGE: <path>
@@ -616,10 +677,10 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|---|---|---|---|
| Technical (perf, CWV, security headers, indexability) | 20% | 30% | |
| Technical (perf, CWV, indexability) | 20% | 30% | |
| On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | |
| SEO Local (NAP, GMB, citations) | 25% | 5% | |
| Off-page (backlinks, mentions, authority) | 10% | 15% | |
| Off-page (unlinked brand mentions — backlinks/authority NOT auditable, §14) | 10% | 15% | |
| Social presence | 10% | 5% | |
| Competitive position | 5% | 10% | |
| Legal compliance | 10% | 5% | |
@@ -628,17 +689,63 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
real users, from STEP 4) when available; otherwise lab PageSpeed
Lighthouse run.
**Security headers are NOT scored here (I4).** `/harden` owns them and
grades them out of 100 with three external validators — pricing them into
this axis too was double-counting the same finding in two reports
(`depth-matrix.md:29` already said drop; this spec contradicted it).
- Dispatched from `/harden` (its prompt says NARROW-SCOPE): headers ARE the
job — audit and score them per its brief, ignore this note.
- Dispatched from `/seo`: do not score CSP, HSTS, X-Frame-Options,
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
cookie flags. STEP 4 still reads them — you need them for the one
carve-out below — but they earn and lose no points here.
**Carve-out — `X-Robots-Tag` stays.** It is an indexing directive wearing a
header's clothes: `noindex` served there deindexes the page as surely as a
meta robots tag. Score it under indexability. That is what
`depth-matrix.md:29` means by "unless it directly affects indexability" —
it is the header that does, and the security headers above are not.
**Drop ≠ silence.** A user who never runs `/harden` must not read a clean
Technical score as clean headers. Whenever depth=FULL, emit in §14:
`Security headers (CSP, HSTS, X-Frame-Options…) — not scored here: /harden
owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden
<url>. Observed live this run: <present list | none observed>.`
Name what you saw. An omission has to stay legible — the same reason
COVERAGE is mandatory in STEP 9.
**Off-page axis note (I1).** Score ONLY the unlinked brand mentions
gathered in STEP 6 (`web_search "<business-name>" -site:<domain>`).
Backlink profile and domain authority have NO data source here — no index,
no API, nothing. NEVER price them into the number: an unmeasured
sub-component cannot be judged, and this axis carries 10-15% of a score
that reaches a client via `/client-handover`. A low mention count is a low
mention count — it is NOT evidence of a weak backlink profile.
Mandatory §14 line whenever depth=FULL, verbatim:
`Backlinks / domain authority — NOT audited: no backlink index wired.
Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs /
Semrush / Majestic. The Off-page score above prices in brand mentions only.`
Weight deliberately unchanged despite the narrower scope: re-deriving it
now, then again when a backlink source lands, would churn historical
scores twice. Revisit the 10/15% only when the axis widens back.
### LOCAL depth — 4 axes
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|---|---|---|---|
| Technical (security headers, indexability, config) | 25% | 35% | |
| Technical (indexability, config) | 25% | 35% | |
| On-page (content, meta, headings, images, video, a11y, i18n) | 35% | 45% | |
| SEO Local (markup, NAP in JSON-LD, legal) | 20% | 5% | |
| Legal compliance (pages, CMP, mentions) | 20% | 15% | |
LOCAL axes not audited (Off-page, Social, Competitive) appear as
`N/A — requires FULL audit` in the report.
`N/A — requires FULL audit` in the report. Off-page is the exception to
that promise: FULL audits its brand-mentions share ONLY — backlinks and
authority are unauditable at EVERY depth (see the Off-page axis note).
Print `N/A — FULL audits brand mentions only` for it, never a bare
"requires FULL audit" that FULL cannot keep.
### Projected code-only score + trajectory to 17/20 (mandatory)
@@ -676,6 +783,8 @@ misroutes the client-handover gate and the user's effort.
```
SEO SCORING (<depth>)
COVERAGE : <N> of <M> sitemap URLs (<P>%) — templates skipped: <list|none>
| <N> pages, total UNKNOWN (no sitemap)
Technical : XX/20 <justification>
On-page : XX/20 <justification>
SEO Local : XX/20 | N/A
@@ -687,6 +796,12 @@ Legal : XX/20 <justification>
SEO GLOBAL (weighted): XX.X/20 (<depth>)
```
**COVERAGE is mandatory, never omitted, never rounded up.** It is the
honesty bound on every page-level axis: On-page and the on-page share of
Technical are extrapolations from the sample. If coverage < 25%, repeat it
in §0 as a major alert — a 17/20 drawn from 3% of a site is not a 17/20, and
`/client-handover` gates on these numbers.
Per user instruction: this score represents **80% of the combined
final score for local B2C (20% for GEO), or 75% for SaaS/national
(25% for GEO)**. The `/seo` dispatcher combines SEO and GEO scores.
+12 -3
View File
@@ -204,9 +204,9 @@ Ask ONCE before dispatching the agents:
```
RAPPORT EXTERNE (optionnel) — un autre regard sur le site :
1. Fichier — déposez l'export (PDF/MD/TXT) dans
`.claude/audits/external/` (ex. `sorank-YYYY-MM-DD.pdf`),
donnez le nom du fichier. (`mkdir -p .claude/audits/external`)
1. Fichier — donnez le chemin de l'export (PDF/MD/TXT), où qu'il soit
(ex. `~/Téléchargements/sorank-2026-07-16.pdf`). Rangement conseillé
mais optionnel : `.claude/audits/external/`.
2. Collé — collez ici le contenu du PDF ou le "prompt pour IA"
que l'outil suggère.
3. Ignorer — continuer sans. Le rapport final recommandera
@@ -348,6 +348,15 @@ audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas,
entity SEO, content shape for AI, AI visibility) — the geo-analyzer
agent runs in parallel and owns those.
Do NOT score security headers either (CSP, HSTS, X-Frame-Options,
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
cookie flags) — `/harden` owns them and grades them 0-100 against three
external validators (`depth-matrix.md:29`). Read them, keep
`X-Robots-Tag` under indexability (it is an indexing directive, not a
security header), and declare the rest in §14 with a "run /harden" pointer
plus what you observed live. Dropping them from the score must not make
them silent.
FILE OWNERSHIP (authoritative, prevents parallel-edit conflicts):
- YOU OWN (read+write): sitemap.xml, image/video sitemaps, .htaccess,
meta tags (title, description, OG, Twitter, canonical, robots meta),