Merge bugfix/seo-geo-integrity-phase1 into develop
This commit is contained in:
@@ -1,5 +1,112 @@
|
|||||||
# TODO
|
# TODO
|
||||||
|
|
||||||
|
## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage)
|
||||||
|
Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0,
|
||||||
|
5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install
|
||||||
|
(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md`
|
||||||
|
deletes our seo-analyzer.md 42K it never installed; extensions/*/install.sh:42
|
||||||
|
wipes settings.json on parse error; skills/seo/SKILL.md:119 injects Skool
|
||||||
|
upsell footer into deliverables). Their code is real (render_page.py 428 l
|
||||||
|
Playwright, url_safety.py 622 l SSRF, 326 tests, 320 pass) — adapt to our
|
||||||
|
fetch.sh contract, do NOT copy wholesale (no fail-open, no tokenstore, no
|
||||||
|
JSON shape).
|
||||||
|
|
||||||
|
Framing: their plus-values map onto OUR integrity gaps — report claims more
|
||||||
|
than it measured. Same bar we held their README to.
|
||||||
|
Seam: `lib/seo-data/fetch.sh` verbs (accounts|crux|queries|inspect|forget)
|
||||||
|
+ fail-open `{"status":"degraded"}` + fixtures + tests. Everything below lands
|
||||||
|
as NEW VERBS. No new architecture.
|
||||||
|
|
||||||
|
### AXE 0 — Integrity (no new deps, hours) — the score currently lies
|
||||||
|
- [ ] I1 Off-page axis scores 10-15% of FULL with ZERO data source (no API,
|
||||||
|
no index) → today fabricated, and it feeds /client-handover. Immediate
|
||||||
|
fix: extend existing LOCAL `N/A — requires FULL audit` pattern to FULL,
|
||||||
|
redistribute weights. Data upgrade later (AXE 3). Honesty now, data after.
|
||||||
|
- [ ] I2 VSI (Visual Stability Index) listed in CWV thresholds but NO path
|
||||||
|
retrieves it — neither CrUX nor PSI expose it. Phantom signal → remove
|
||||||
|
or source.
|
||||||
|
- [ ] I3 **SAFETY** /geo standalone: geo/SKILL.md (125 l) has no STEP 0, no
|
||||||
|
confirmed-NAP collection — but geo-analyzer OWNS JSON-LD NAP. Standalone
|
||||||
|
/geo on a local business can write unverified NAP with zero LRN-032
|
||||||
|
protection. Real bug, not cosmetic.
|
||||||
|
- [ ] I4 Security headers counted 3× (seo-analyzer STEP 4 scores them in
|
||||||
|
Technical axis; depth-matrix.md says drop unless indexability; /harden
|
||||||
|
re-audits /100 with 3 validators). Contradiction between dedup rule and
|
||||||
|
agent spec → pick one owner.
|
||||||
|
- [ ] I5 Report says "audit", measured 5-15 sampled pages. State coverage %
|
||||||
|
explicitly in §0 until AXE 2 lands.
|
||||||
|
|
||||||
|
### AXE 1 — Free wins on auth we ALREADY have (fetch.sh verbs)
|
||||||
|
- [ ] W1 `richresults` verb — GSC URL Inspection already returns
|
||||||
|
`richResultsResult`; our OAuth already carries the scope. Programmatic
|
||||||
|
rich-results validation on real Google data. **BEATS claude-seo**: their
|
||||||
|
README:314 "dual validator (Rich Results Test + Markup Validator)" is
|
||||||
|
FALSE — grep of all .py = zero calls, they are hyperlinks a human clicks.
|
||||||
|
Today our JSON-LD validity is LLM-read only.
|
||||||
|
- [ ] W2 `bing` verb — Bing Webmaster API, free. Closes the Google/Bing
|
||||||
|
asymmetry (Google = full OAuth layer, Bing = manual checklist) while
|
||||||
|
/geo targets ChatGPT Search, which indexes via Bing. Strategic, not cosmetic.
|
||||||
|
- [ ] W3 `sameas` resolution check — trivial curl loop. entity-seo.md lists
|
||||||
|
"sameAs pointing to dead profiles" as a known error class and never
|
||||||
|
checks it. ~10 lines.
|
||||||
|
|
||||||
|
### AXE 2 — Coverage (biggest lever: ~97% of a 500-page site unseen today)
|
||||||
|
- [ ] C1 `crawl` verb — sitemap-driven URL discovery (we ALREADY fetch
|
||||||
|
sitemap.xml) + deterministic sampling + coverage % reported. No Chromium,
|
||||||
|
no paid API. Turns "5-15 LLM-chosen pages" into measured coverage.
|
||||||
|
Tradeoff vs claude-seo's link-following 500-page crawl: cheaper, but
|
||||||
|
misses unlinked/unsitemapped pages — accept + disclose.
|
||||||
|
- [ ] C2 Dupe/cannibalization detection — becomes possible once N pages in
|
||||||
|
hand: compare titles/H1/canonicals across the set. Free, unblocked by C1.
|
||||||
|
- [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated
|
||||||
|
as checks with no command to compute them. C1 unblocks real computation.
|
||||||
|
|
||||||
|
### AXE 3 — Off-page real (upgrades I1)
|
||||||
|
- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph
|
||||||
|
(data.commoncrawl.org/projects/hyperlinkgraph), free, no key.
|
||||||
|
- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap
|
||||||
|
health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine
|
||||||
|
exactly.
|
||||||
|
- [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth
|
||||||
|
already there" — I doubt it: Search Console API v3 has no links endpoint
|
||||||
|
(links report is UI-only AFAIK). Verify before planning on it. Do not
|
||||||
|
assert.
|
||||||
|
|
||||||
|
### AXE 4 — SPA blindness (dep decision — needs arbitrage)
|
||||||
|
- [ ] R1 `render` verb — Playwright, GATED on SPA detection (STEP 2 already
|
||||||
|
detects framework + rendering mode). Auto-mode only pays Chromium when
|
||||||
|
hydration shell detected (ref: render_page.py:226 logic, adapt not copy).
|
||||||
|
- [ ] R2 ARBITRAGE: heavy dep (Chromium ~300MB) vs our bash+curl purity.
|
||||||
|
Cheaper honest alternative: on SPA, REFUSE to score on-page rather than
|
||||||
|
score it wrong (today: curl reads source, not hydrated DOM → every
|
||||||
|
meta/JSON-LD/heading/img grep is blind, compensated only by a §0 flag).
|
||||||
|
|
||||||
|
### AXE 5 — Hardening + regression (lower priority)
|
||||||
|
- [ ] H1 SSRF guard on curl paths — both agents curl user-supplied domains.
|
||||||
|
Our own CLAUDE.md doctrine says "never trust user input". url_safety.py
|
||||||
|
(622 l, obfuscated-IPv4 decode, DNS pinning) is a solid reference.
|
||||||
|
- [ ] H2 `drift` baseline (SQLite) — SEO.md Historique keeps only date+score+
|
||||||
|
key changes. Their seo-drift is on-page regression detection, NOT rank
|
||||||
|
tracking (common misread). Optional.
|
||||||
|
|
||||||
|
### NOT DOING (explicit, with reason)
|
||||||
|
- Keyword volumes → Google Ads Tier 3 needs ACTIVE ad spend (~$150-300/mo);
|
||||||
|
without spend the API returns buckets ("1K-10K"). Their own detect_tier()
|
||||||
|
never even returns 3 (google_auth.py:642-724 caps at 2) + google-ads absent
|
||||||
|
from requirements.txt. Not worth it.
|
||||||
|
- Real AI SoV (ChatGPT/Perplexity citation tracking) → paid everywhere
|
||||||
|
(SE Ranking/Profound/DataForSEO). Our current honest "not testable, here's
|
||||||
|
what we measured instead" disclosure BEATS faking it. Keep.
|
||||||
|
- Installing the plugin / +33 skills namespace → see destructive paths above.
|
||||||
|
|
||||||
|
### Keep (already beats claude-seo — do not regress)
|
||||||
|
FR legal (LCEN/RGPD-ePrivacy/DGCCRF L121-1 — their whole repo: 2 hits, and
|
||||||
|
dma-consent-mode-v2.md:27 tells the agent to stay out) · fix-bundle +
|
||||||
|
ownership matrix + serial apply (their 18 agents are report-only, no
|
||||||
|
ownership discipline) · trajectory-to-17/20 + honest code ceiling (theirs is
|
||||||
|
flat 0-100, no legal axis) · llms.txt honest framing · NAP anti-dup-seed
|
||||||
|
(LRN-032).
|
||||||
|
|
||||||
## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068)
|
## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068)
|
||||||
- [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop
|
- [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop
|
||||||
- [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail
|
- [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail
|
||||||
|
|||||||
+57
-7
@@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the
|
|||||||
|
|
||||||
## Context — why GEO is its own discipline in 2026
|
## Context — why GEO is its own discipline in 2026
|
||||||
|
|
||||||
- AI Overviews trigger on ~48% of Google searches (April 2026).
|
- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google
|
||||||
- ChatGPT processes 2.5B queries/day.
|
searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner
|
||||||
- Gartner projects commercial organic search traffic to fall 25% by
|
projects commercial organic search traffic to fall 25% by end-2026 as
|
||||||
end-2026 as discovery shifts to AI engines.
|
discovery shifts to AI engines. Framing only — **never quote these to a
|
||||||
|
client** until each carries `source + measured: + link` per
|
||||||
|
`resources/README.md`. GEO is worth doing on mechanism; it does not need
|
||||||
|
these numbers to be true.
|
||||||
- Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org)
|
- Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org)
|
||||||
but the optimization levers differ: entity clarity, definition
|
but the optimization levers differ: entity clarity, definition
|
||||||
architecture, citable stats, crawler permissions.
|
architecture, citable stats, crawler permissions.
|
||||||
@@ -141,6 +144,16 @@ If called standalone via `/geo`, gather:
|
|||||||
|
|
||||||
## STEP 2 — DETECT CONTEXT `[both]`
|
## STEP 2 — DETECT CONTEXT `[both]`
|
||||||
|
|
||||||
|
**FIRST — the CWD must BE the audited site.** You grep the current working
|
||||||
|
directory; no dispatcher checks that it matches the target domain. If a URL
|
||||||
|
was supplied and the CWD shows no web project at all (no `package.json` /
|
||||||
|
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
|
||||||
|
signals contradict the domain, STOP and report:
|
||||||
|
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
|
||||||
|
confirm live-only audit (LOCAL findings will be N/A).`
|
||||||
|
Never grep one codebase while curling another: the live half looks right,
|
||||||
|
the code half is fiction, and the report reads as authoritative.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Framework (reuse detection from seo-analyzer if available)
|
# Framework (reuse detection from seo-analyzer if available)
|
||||||
ls package.json composer.json Gemfile Cargo.toml go.mod 2>/dev/null
|
ls package.json composer.json Gemfile Cargo.toml go.mod 2>/dev/null
|
||||||
@@ -360,7 +373,9 @@ action (G5 batch, confirmation needed — visible page creation).
|
|||||||
|
|
||||||
**Local business:**
|
**Local business:**
|
||||||
- [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.)
|
- [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.)
|
||||||
- [ ] NAP consistent with GMB
|
- [ ] NAP consistent with GMB — **direction rule applies** (Data integrity:
|
||||||
|
never pick a value from source majority; no canonical → no directional
|
||||||
|
fix)
|
||||||
- [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable
|
- [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable
|
||||||
- [ ] `areaServed` lists served cities/regions
|
- [ ] `areaServed` lists served cities/regions
|
||||||
- [ ] `openingHoursSpecification` matches reality
|
- [ ] `openingHoursSpecification` matches reality
|
||||||
@@ -447,6 +462,12 @@ Load: `~/.claude/agents/resources/content-shape-for-ai.md`
|
|||||||
|
|
||||||
Sample 5-10 key pages (homepage + top service/blog pages). For each:
|
Sample 5-10 key pages (homepage + top service/blog pages). For each:
|
||||||
|
|
||||||
|
**Record the denominator.** This samples; the report says "audit". Count the
|
||||||
|
URLs in `sitemap.xml` for the coverage ratio, and carry it into the GEO
|
||||||
|
SCORING block. No sitemap → total UNKNOWN, say so. Content shape is the
|
||||||
|
axis most damaged by silent sampling: it is judged per page, so a 6-page
|
||||||
|
sample of a 300-page site says nothing about the other 294.
|
||||||
|
|
||||||
### Checks
|
### Checks
|
||||||
|
|
||||||
1. **Definition Lead** — does the first sentence (or H1) follow
|
1. **Definition Lead** — does the first sentence (or H1) follow
|
||||||
@@ -574,6 +595,7 @@ Score each axis. Use concrete findings from STEP 2-9.
|
|||||||
|
|
||||||
```
|
```
|
||||||
GEO SCORING (<depth>)
|
GEO SCORING (<depth>)
|
||||||
|
COVERAGE : <N> of <M> sitemap URLs (<P>%) | <N> pages, total UNKNOWN
|
||||||
AI Crawlers Policy : XX/20 <justification>
|
AI Crawlers Policy : XX/20 <justification>
|
||||||
llms.txt : XX/20 <justification>
|
llms.txt : XX/20 <justification>
|
||||||
Schema.org for AI : XX/20 <justification>
|
Schema.org for AI : XX/20 <justification>
|
||||||
@@ -584,6 +606,12 @@ AI Visibility (live) : XX/20 | N/A (LOCAL)
|
|||||||
GEO GLOBAL (weighted) : XX.X/20 (<depth>)
|
GEO GLOBAL (weighted) : XX.X/20 (<depth>)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**COVERAGE is mandatory, never omitted, never rounded up.** It bounds the
|
||||||
|
per-page axes — Content Shape above all, and the page-level share of
|
||||||
|
Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected:
|
||||||
|
robots.txt and llms.txt are single files, fully read. Say which is which
|
||||||
|
rather than letting one ratio discredit the whole report.
|
||||||
|
|
||||||
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
||||||
local, 25% for national/SaaS/content.**
|
local, 25% for national/SaaS/content.**
|
||||||
|
|
||||||
@@ -895,9 +923,31 @@ PROCHAINE ETAPE : <highest-priority>
|
|||||||
- **No invented entity data.** Never write a fake Wikidata QID, fake
|
- **No invented entity data.** Never write a fake Wikidata QID, fake
|
||||||
`sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown →
|
`sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown →
|
||||||
placeholder `[À COMPLÉTER]` or omit.
|
placeholder `[À COMPLÉTER]` or omit.
|
||||||
|
- **NAP direction rule (LRN-032).** You own JSON-LD NAP, so this binds you
|
||||||
|
whoever called you — `/seo` passes a canonical, standalone `/geo` does
|
||||||
|
not. NEVER infer a correct NAP value from source majority: on-site
|
||||||
|
sources (JSON-LD, footer, settings DB, legal pages) usually descend from
|
||||||
|
ONE seed and can all carry the same wrong value — the single diverging
|
||||||
|
source may be the only one a human actually corrected. Direction of fix:
|
||||||
|
- Diverging from a CONFIRMED canonical field (passed by `/seo` STEP 0)
|
||||||
|
→ fix the diverging source.
|
||||||
|
- Canonical UNCONFIRMED or absent (the standalone `/geo` case) → report
|
||||||
|
the divergence WITHOUT a directional fix; escalate as a user question
|
||||||
|
("which value is correct?") in §11.
|
||||||
|
No G2/G6 item may write or rewrite a NAP value that no confirmed
|
||||||
|
canonical backs — **creating** a `LocalBusiness` from scratch included:
|
||||||
|
unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling
|
||||||
|
on-site source.
|
||||||
- **Remove deprecated schemas rather than keep broken ones.**
|
- **Remove deprecated schemas rather than keep broken ones.**
|
||||||
- **Cite sources.** When emitting stats in the report, link
|
- **Cite sources, and only citable ones.** A stat reaches the client only
|
||||||
`content-shape-for-ai.md` research citations.
|
if it carries `source + measured: + link` per `resources/README.md`.
|
||||||
|
Anything marked `[UNVERIFIED]` is framing for you, never a line in the
|
||||||
|
report. Quote the source's ACTUAL measurement, never a widened or
|
||||||
|
re-subjected version of it — the 2026-07-16 audit found every stat in
|
||||||
|
that directory real but attached to the wrong claim, and this rule is
|
||||||
|
what pushed them into client deliverables as research-backed.
|
||||||
|
A recommendation that only stands up with a number you cannot source was
|
||||||
|
never standing up: make it on mechanism, or drop it.
|
||||||
|
|
||||||
### Process
|
### Process
|
||||||
- **Every user action lists automation options.** Mandatory from
|
- **Every user action lists automation options.** Mandatory from
|
||||||
|
|||||||
@@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current.
|
|||||||
|
|
||||||
These files capture state as of 2026-04. Crawler lists, Schema.org
|
These files capture state as of 2026-04. Crawler lists, Schema.org
|
||||||
deprecations, and tool landscape shift fast. Agents MUST cross-check
|
deprecations, and tool landscape shift fast. Agents MUST cross-check
|
||||||
via WebSearch on each run when FULL depth is selected.
|
crawler lists and tool names via WebSearch on each run when FULL depth is
|
||||||
|
selected.
|
||||||
|
|
||||||
|
## Citation standard (mandatory for every statistic)
|
||||||
|
|
||||||
|
**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO
|
||||||
|
blogs cross-cite each other into a consensus that looks like corroboration.
|
||||||
|
Two 2026-07-16 audits of this directory show how it fails:
|
||||||
|
|
||||||
|
- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in
|
||||||
|
`seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already
|
||||||
|
collected it. It is absent from the CrUX API metric list and from
|
||||||
|
web.dev. WebSearch returned the echo, not the truth.
|
||||||
|
- Every stat in this directory was real **and attached to the wrong
|
||||||
|
subject**: the GEO paper's 40% (all methods) pinned on one technique;
|
||||||
|
LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay;
|
||||||
|
AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with
|
||||||
|
its meaning inverted; a smart-speaker adoption figure sold as voice-search
|
||||||
|
share.
|
||||||
|
|
||||||
|
The failure mode is not invention — it is **plausible recombination**, which
|
||||||
|
is exactly what a model half-remembering a search result produces. So the
|
||||||
|
format has to make an unsourced number conspicuous:
|
||||||
|
|
||||||
|
```
|
||||||
|
<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
|
||||||
|
measured> — <link>
|
||||||
|
```
|
||||||
|
|
||||||
|
`measured:` is the field that catches it. All four errors above survive a
|
||||||
|
source name; none survives having to state the source's real measurement
|
||||||
|
next to the claim.
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
1. **Primary source or no number.** Peer-reviewed paper, the vendor's own
|
||||||
|
published study, or an official API/doc. `developer.chrome.com/docs/crux`
|
||||||
|
is decisive for metrics: what CrUX cannot return, we cannot score.
|
||||||
|
2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast,
|
||||||
|
Ahrefs publish useful data and sell products — say "vendor".
|
||||||
|
3. **Never widen scope.** An aggregate result is not a per-technique result.
|
||||||
|
4. **No number beats a wrong number.** A recommendation that only stands up
|
||||||
|
with a fabricated statistic was never standing up. Delete the stat, keep
|
||||||
|
the recommendation if it survives on mechanism.
|
||||||
|
5. **Unverified ⇒ labelled.** `[UNVERIFIED — <date>]` inline. Never quote an
|
||||||
|
unverified number to a client: `geo-analyzer.md` ("Cite sources") sends
|
||||||
|
these into client reports as research-backed.
|
||||||
|
|
||||||
## Loading pattern
|
## Loading pattern
|
||||||
|
|
||||||
|
|||||||
@@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers
|
|||||||
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
|
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
|
||||||
Overviews.
|
Overviews.
|
||||||
|
|
||||||
Context: Google AI Overviews trigger on ~48% of searches; ChatGPT
|
Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of
|
||||||
processes 2.5B queries/day; Gartner projects commercial organic
|
searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial
|
||||||
search traffic will drop 25% by 2026. Monitoring is no longer optional.
|
organic search traffic will drop 25% by 2026.
|
||||||
|
|
||||||
|
> Not checked against primary sources in the 2026-07-16 audit that corrected
|
||||||
|
> the rest of this directory — flagged rather than asserted or deleted, per
|
||||||
|
> the citation standard in `README.md` (rule 5). The Gartner projection at
|
||||||
|
> least names its source; the other two float. Treat all three as
|
||||||
|
> motivation, not evidence: **do NOT quote them to a client** until each
|
||||||
|
> carries `source + measured: + link`. Their only job here is to explain why
|
||||||
|
> this file exists, and that argument does not need numbers.
|
||||||
|
|
||||||
## Commercial tools
|
## Commercial tools
|
||||||
|
|
||||||
|
|||||||
@@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density.
|
|||||||
|
|
||||||
### 4. Citations and statistics (strongest measured lever)
|
### 4. Citations and statistics (strongest measured lever)
|
||||||
|
|
||||||
Adding peer-cited statistics with clear sources increases AI visibility
|
Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024)
|
||||||
**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine
|
report that their optimisation methods **collectively** boost visibility
|
||||||
Optimization").
|
**by up to 40%** in generative-engine responses, and state the effect
|
||||||
|
**varies across domains**. Citations/statistics/quotations are among those
|
||||||
|
methods.
|
||||||
|
|
||||||
|
> **Attribute this correctly.** Until 2026-07-16 this section read "Adding
|
||||||
|
> peer-cited statistics with clear sources increases AI visibility by up to
|
||||||
|
> 40%" — pinning the paper's *aggregate* result on this *one* technique. The
|
||||||
|
> paper publishes no separate figure per technique. When quoting it to a
|
||||||
|
> client: "up to 40%, across the method set, domain-dependent" — never "+40%
|
||||||
|
> if you add stats".
|
||||||
|
|
||||||
Pattern: embed specific numbers with attribution.
|
Pattern: embed specific numbers with attribution.
|
||||||
|
|
||||||
@@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure:
|
|||||||
|
|
||||||
### 6. Freshness signals
|
### 6. Freshness signals
|
||||||
|
|
||||||
Pages not updated at least quarterly are **3x more likely to lose AI
|
Freshness is a real retrieval input: RAG systems fetch live and read
|
||||||
citations** (LLMRefs 2026 study).
|
timestamps, so a page updated this quarter carries a stronger recency
|
||||||
|
signal than the same page last touched years ago. LLMrefs (a **vendor**,
|
||||||
|
not peer review) reports cited content running **~25.7% fresher** than
|
||||||
|
organic top-10 across ~17M citations. Substantive updates only — bumping a
|
||||||
|
date string is not freshness.
|
||||||
|
|
||||||
|
> **The "3x" that lived here was grafted from another claim.** Until
|
||||||
|
> 2026-07-16 this read "Pages not updated at least quarterly are 3x more
|
||||||
|
> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x"
|
||||||
|
> says **brand mentions correlate ~3x more strongly with AI visibility than
|
||||||
|
> backlinks** — a different subject entirely. No source supports a quarterly
|
||||||
|
> decay multiplier. Recommend quarterly refresh on its merits; do not price
|
||||||
|
> it with a borrowed number.
|
||||||
|
|
||||||
What to maintain:
|
What to maintain:
|
||||||
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
|
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
|
||||||
|
|||||||
@@ -21,8 +21,20 @@ existing instances. They no longer produce rich results.
|
|||||||
|
|
||||||
### QAPage — single Q&A format
|
### QAPage — single Q&A format
|
||||||
|
|
||||||
Pages cited 58% more often by ChatGPT vs basic Article schema.
|
Use when the page is built around ONE primary question. Emitting the type
|
||||||
Use when the page is built around ONE primary question.
|
that matches the content shape beats wrapping everything in a generic
|
||||||
|
`Article`.
|
||||||
|
|
||||||
|
> **No lift figure here — the one that lived here was wrong.** Until
|
||||||
|
> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic
|
||||||
|
> Article schema", uncited. Nothing supports it. The nearest real number is
|
||||||
|
> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews /
|
||||||
|
> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%**
|
||||||
|
> of cited sources — a *prevalence* count for a *different type* — while
|
||||||
|
> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim
|
||||||
|
> it was propping up. Q&A shape is still worth doing on genuinely
|
||||||
|
> single-question pages; it is not worth a fabricated number. Do NOT quote a
|
||||||
|
> QAPage lift % to a client — there isn't one.
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -81,8 +93,16 @@ visible content.
|
|||||||
|
|
||||||
### Speakable — voice + AI extraction marker
|
### Speakable — voice + AI extraction marker
|
||||||
|
|
||||||
62% of searches in 2026 involve voice. Speakable flags the passage
|
Speakable flags the passage best suited for voice readout and AI summary.
|
||||||
best suited for voice readout and AI summary.
|
|
||||||
|
> **No voice-share figure — the one that lived here was a conflation.**
|
||||||
|
> Until 2026-07-16 this read "62% of searches in 2026 involve voice",
|
||||||
|
> uncited. No primary source carries it; 62% circulates as a *smart-speaker
|
||||||
|
> adoption* number, not a share of searches. It is the same family as the
|
||||||
|
> "50% of searches will be voice by 2020" myth — attributed to ComScore,
|
||||||
|
> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable
|
||||||
|
> is cheap and harmless, so keep recommending it on TL;DR / summary blocks —
|
||||||
|
> but justify it by extraction shape, never by a voice-share statistic.
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
|
|||||||
+121
-6
@@ -81,6 +81,17 @@ hreflang, infer from detected URL structures.
|
|||||||
|
|
||||||
## STEP 2 — DETECT TECHNICAL CONTEXT `[both]`
|
## STEP 2 — DETECT TECHNICAL CONTEXT `[both]`
|
||||||
|
|
||||||
|
**FIRST — the CWD must BE the audited site.** You grep the current working
|
||||||
|
directory; no dispatcher checks that it matches TARGET_URL. If a URL was
|
||||||
|
supplied and the CWD shows no web project at all (no `package.json` /
|
||||||
|
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
|
||||||
|
signals contradict the domain, STOP and report:
|
||||||
|
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
|
||||||
|
confirm live-only audit (LOCAL findings will be N/A).`
|
||||||
|
Never grep one codebase while curling another: the live half looks right,
|
||||||
|
the code half is fiction, and the report reads as authoritative. `/harden`
|
||||||
|
inherits this agent for its config axis, so the mismatch propagates there.
|
||||||
|
|
||||||
### Framework & rendering
|
### Framework & rendering
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -148,6 +159,23 @@ RECOMMENDATION : KEEP & CONFIGURE plugin | INSTALL <plugin> (P0 quick win) | M
|
|||||||
|
|
||||||
### Infrastructure signals
|
### Infrastructure signals
|
||||||
|
|
||||||
|
**Origin vs edge — never infer the stack from `server:`.** That header names
|
||||||
|
whatever answered: usually the EDGE (Cloudflare, Scaleway/OVH front, CDN,
|
||||||
|
load balancer), not the origin. Apache behind an nginx front is a standard
|
||||||
|
topology — TLS terminated upstream, the origin sees plain HTTP plus
|
||||||
|
`X-Forwarded-Proto`.
|
||||||
|
- Repo `.htaccess` + `server: nginx` = NOT drift, NOT dead config. Do not
|
||||||
|
flag it, do not propose migrating it.
|
||||||
|
- Never move headers into an `nginx.conf` absent from the repo. Server-side
|
||||||
|
config you cannot read is a §14 gap, not a finding.
|
||||||
|
- A header present live but in no repo config = "set upstream", never
|
||||||
|
"missing".
|
||||||
|
|
||||||
|
`/harden` reuses this agent for its entire config-hardening axis, so a wrong
|
||||||
|
topology call scores a client's server config against a file that never ran.
|
||||||
|
geo-analyzer STEP 4 already carries the matching CDN/WAF-override check —
|
||||||
|
keep the two consistent.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Server / hosting
|
# Server / hosting
|
||||||
ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null
|
ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null
|
||||||
@@ -216,6 +244,12 @@ anonymous PageSpeed lab data and STEP 4/STEP 11 emit the §11 user action
|
|||||||
|
|
||||||
### HTTP headers & security
|
### HTTP headers & security
|
||||||
|
|
||||||
|
**Read them; score them only for `/harden` (I4).** This section stays — the
|
||||||
|
raw headers are needed for `X-Robots-Tag`, canonical/redirect coherence, and
|
||||||
|
the §14 observed-list. But under `/seo` the security headers themselves are
|
||||||
|
out of scope for scoring: see the Technical axis note in STEP 9. Under
|
||||||
|
`/harden` they are the entire job. Reading is not scoring.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
DOMAIN="<production-domain>"
|
DOMAIN="<production-domain>"
|
||||||
|
|
||||||
@@ -247,8 +281,21 @@ Evaluate each present/missing:
|
|||||||
- **LCP** (Largest Contentful Paint) — < 2.5s
|
- **LCP** (Largest Contentful Paint) — < 2.5s
|
||||||
- **INP** (Interaction to Next Paint) — < 200ms (replaced FID in Mar 2024)
|
- **INP** (Interaction to Next Paint) — < 200ms (replaced FID in Mar 2024)
|
||||||
- **CLS** (Cumulative Layout Shift) — < 0.1
|
- **CLS** (Cumulative Layout Shift) — < 0.1
|
||||||
- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web
|
|
||||||
Vitals 2.0
|
**Core Web Vitals are exactly these three** (web.dev/articles/vitals,
|
||||||
|
verified 2026-07-16). Google ships threshold changes with prior notice on a
|
||||||
|
predictable annual cadence — a "new CWV" that only SEO blogs know about does
|
||||||
|
not exist. Before adding a metric here, confirm it against a PRIMARY source:
|
||||||
|
web.dev, the Chromium blog, or `developer.chrome.com/docs/crux/api` — that
|
||||||
|
API metric list is decisive, because a metric CrUX cannot return is a metric
|
||||||
|
we cannot score.
|
||||||
|
|
||||||
|
**WebSearch is not confirmation.** SEO blogs cross-cite each other into fake
|
||||||
|
consensus. A "VSI (Visual Stability Index) — new 2026 signal, Core Web
|
||||||
|
Vitals 2.0" line lived here until 2026-07-16 on exactly that basis: ten
|
||||||
|
blogs asserted it, several claimed CrUX was already collecting it, and it is
|
||||||
|
absent from both the CrUX API metric list and web.dev. Stated as fact, in a
|
||||||
|
threshold list, in client-facing audits.
|
||||||
|
|
||||||
When a GSC account+property were passed in context, fetch CrUX field
|
When a GSC account+property were passed in context, fetch CrUX field
|
||||||
data first (**tilde path mandatory** — this agent runs from the
|
data first (**tilde path mandatory** — this agent runs from the
|
||||||
@@ -359,8 +406,22 @@ Fetch rendered HTML. Extract and analyze:
|
|||||||
|
|
||||||
## STEP 5 — ON-PAGE AUDIT `[both]`
|
## STEP 5 — ON-PAGE AUDIT `[both]`
|
||||||
|
|
||||||
|
**Record the denominator BEFORE sampling.** This step samples; the report
|
||||||
|
says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the
|
||||||
|
`head -50` in STEP 4 is a preview, not a count). That count is the coverage
|
||||||
|
denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap
|
||||||
|
→ denominator unknown: say so, never let silence imply full coverage. On a
|
||||||
|
500-page site a 12-page sample is 2.4% — the On-page score is an
|
||||||
|
extrapolation from it, and the reader cannot know that unless you print it.
|
||||||
|
|
||||||
### Meta tags per page (sample 5-15 key pages)
|
### Meta tags per page (sample 5-15 key pages)
|
||||||
|
|
||||||
|
Sample by risk, not convenience: homepage + top templates (one per page
|
||||||
|
type: service, city, blog, product, legal) + any page GSC flags as a
|
||||||
|
position 4-10 quick win. Same template audited twice buys nothing; an
|
||||||
|
un-sampled template is an un-audited template — name the templates you
|
||||||
|
skipped.
|
||||||
|
|
||||||
For each sampled page:
|
For each sampled page:
|
||||||
```
|
```
|
||||||
PAGE: <path>
|
PAGE: <path>
|
||||||
@@ -616,10 +677,10 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
|||||||
|
|
||||||
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| Technical (perf, CWV, security headers, indexability) | 20% | 30% | |
|
| Technical (perf, CWV, indexability) | 20% | 30% | |
|
||||||
| On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | |
|
| On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | |
|
||||||
| SEO Local (NAP, GMB, citations) | 25% | 5% | |
|
| SEO Local (NAP, GMB, citations) | 25% | 5% | |
|
||||||
| Off-page (backlinks, mentions, authority) | 10% | 15% | |
|
| Off-page (unlinked brand mentions — backlinks/authority NOT auditable, §14) | 10% | 15% | |
|
||||||
| Social presence | 10% | 5% | |
|
| Social presence | 10% | 5% | |
|
||||||
| Competitive position | 5% | 10% | |
|
| Competitive position | 5% | 10% | |
|
||||||
| Legal compliance | 10% | 5% | |
|
| Legal compliance | 10% | 5% | |
|
||||||
@@ -628,17 +689,63 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
|||||||
real users, from STEP 4) when available; otherwise lab PageSpeed
|
real users, from STEP 4) when available; otherwise lab PageSpeed
|
||||||
Lighthouse run.
|
Lighthouse run.
|
||||||
|
|
||||||
|
**Security headers are NOT scored here (I4).** `/harden` owns them and
|
||||||
|
grades them out of 100 with three external validators — pricing them into
|
||||||
|
this axis too was double-counting the same finding in two reports
|
||||||
|
(`depth-matrix.md:29` already said drop; this spec contradicted it).
|
||||||
|
- Dispatched from `/harden` (its prompt says NARROW-SCOPE): headers ARE the
|
||||||
|
job — audit and score them per its brief, ignore this note.
|
||||||
|
- Dispatched from `/seo`: do not score CSP, HSTS, X-Frame-Options,
|
||||||
|
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
|
||||||
|
cookie flags. STEP 4 still reads them — you need them for the one
|
||||||
|
carve-out below — but they earn and lose no points here.
|
||||||
|
|
||||||
|
**Carve-out — `X-Robots-Tag` stays.** It is an indexing directive wearing a
|
||||||
|
header's clothes: `noindex` served there deindexes the page as surely as a
|
||||||
|
meta robots tag. Score it under indexability. That is what
|
||||||
|
`depth-matrix.md:29` means by "unless it directly affects indexability" —
|
||||||
|
it is the header that does, and the security headers above are not.
|
||||||
|
|
||||||
|
**Drop ≠ silence.** A user who never runs `/harden` must not read a clean
|
||||||
|
Technical score as clean headers. Whenever depth=FULL, emit in §14:
|
||||||
|
`Security headers (CSP, HSTS, X-Frame-Options…) — not scored here: /harden
|
||||||
|
owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden
|
||||||
|
<url>. Observed live this run: <present list | none observed>.`
|
||||||
|
Name what you saw. An omission has to stay legible — the same reason
|
||||||
|
COVERAGE is mandatory in STEP 9.
|
||||||
|
|
||||||
|
**Off-page axis note (I1).** Score ONLY the unlinked brand mentions
|
||||||
|
gathered in STEP 6 (`web_search "<business-name>" -site:<domain>`).
|
||||||
|
Backlink profile and domain authority have NO data source here — no index,
|
||||||
|
no API, nothing. NEVER price them into the number: an unmeasured
|
||||||
|
sub-component cannot be judged, and this axis carries 10-15% of a score
|
||||||
|
that reaches a client via `/client-handover`. A low mention count is a low
|
||||||
|
mention count — it is NOT evidence of a weak backlink profile.
|
||||||
|
|
||||||
|
Mandatory §14 line whenever depth=FULL, verbatim:
|
||||||
|
`Backlinks / domain authority — NOT audited: no backlink index wired.
|
||||||
|
Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs /
|
||||||
|
Semrush / Majestic. The Off-page score above prices in brand mentions only.`
|
||||||
|
|
||||||
|
Weight deliberately unchanged despite the narrower scope: re-deriving it
|
||||||
|
now, then again when a backlink source lands, would churn historical
|
||||||
|
scores twice. Revisit the 10/15% only when the axis widens back.
|
||||||
|
|
||||||
### LOCAL depth — 4 axes
|
### LOCAL depth — 4 axes
|
||||||
|
|
||||||
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| Technical (security headers, indexability, config) | 25% | 35% | |
|
| Technical (indexability, config) | 25% | 35% | |
|
||||||
| On-page (content, meta, headings, images, video, a11y, i18n) | 35% | 45% | |
|
| On-page (content, meta, headings, images, video, a11y, i18n) | 35% | 45% | |
|
||||||
| SEO Local (markup, NAP in JSON-LD, legal) | 20% | 5% | |
|
| SEO Local (markup, NAP in JSON-LD, legal) | 20% | 5% | |
|
||||||
| Legal compliance (pages, CMP, mentions) | 20% | 15% | |
|
| Legal compliance (pages, CMP, mentions) | 20% | 15% | |
|
||||||
|
|
||||||
LOCAL axes not audited (Off-page, Social, Competitive) appear as
|
LOCAL axes not audited (Off-page, Social, Competitive) appear as
|
||||||
`N/A — requires FULL audit` in the report.
|
`N/A — requires FULL audit` in the report. Off-page is the exception to
|
||||||
|
that promise: FULL audits its brand-mentions share ONLY — backlinks and
|
||||||
|
authority are unauditable at EVERY depth (see the Off-page axis note).
|
||||||
|
Print `N/A — FULL audits brand mentions only` for it, never a bare
|
||||||
|
"requires FULL audit" that FULL cannot keep.
|
||||||
|
|
||||||
### Projected code-only score + trajectory to 17/20 (mandatory)
|
### Projected code-only score + trajectory to 17/20 (mandatory)
|
||||||
|
|
||||||
@@ -676,6 +783,8 @@ misroutes the client-handover gate and the user's effort.
|
|||||||
|
|
||||||
```
|
```
|
||||||
SEO SCORING (<depth>)
|
SEO SCORING (<depth>)
|
||||||
|
COVERAGE : <N> of <M> sitemap URLs (<P>%) — templates skipped: <list|none>
|
||||||
|
| <N> pages, total UNKNOWN (no sitemap)
|
||||||
Technical : XX/20 <justification>
|
Technical : XX/20 <justification>
|
||||||
On-page : XX/20 <justification>
|
On-page : XX/20 <justification>
|
||||||
SEO Local : XX/20 | N/A
|
SEO Local : XX/20 | N/A
|
||||||
@@ -687,6 +796,12 @@ Legal : XX/20 <justification>
|
|||||||
SEO GLOBAL (weighted): XX.X/20 (<depth>)
|
SEO GLOBAL (weighted): XX.X/20 (<depth>)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**COVERAGE is mandatory, never omitted, never rounded up.** It is the
|
||||||
|
honesty bound on every page-level axis: On-page and the on-page share of
|
||||||
|
Technical are extrapolations from the sample. If coverage < 25%, repeat it
|
||||||
|
in §0 as a major alert — a 17/20 drawn from 3% of a site is not a 17/20, and
|
||||||
|
`/client-handover` gates on these numbers.
|
||||||
|
|
||||||
Per user instruction: this score represents **80% of the combined
|
Per user instruction: this score represents **80% of the combined
|
||||||
final score for local B2C (20% for GEO), or 75% for SaaS/national
|
final score for local B2C (20% for GEO), or 75% for SaaS/national
|
||||||
(25% for GEO)**. The `/seo` dispatcher combines SEO and GEO scores.
|
(25% for GEO)**. The `/seo` dispatcher combines SEO and GEO scores.
|
||||||
|
|||||||
+12
-3
@@ -204,9 +204,9 @@ Ask ONCE before dispatching the agents:
|
|||||||
```
|
```
|
||||||
RAPPORT EXTERNE (optionnel) — un autre regard sur le site :
|
RAPPORT EXTERNE (optionnel) — un autre regard sur le site :
|
||||||
|
|
||||||
1. Fichier — déposez l'export (PDF/MD/TXT) dans
|
1. Fichier — donnez le chemin de l'export (PDF/MD/TXT), où qu'il soit
|
||||||
`.claude/audits/external/` (ex. `sorank-YYYY-MM-DD.pdf`),
|
(ex. `~/Téléchargements/sorank-2026-07-16.pdf`). Rangement conseillé
|
||||||
donnez le nom du fichier. (`mkdir -p .claude/audits/external`)
|
mais optionnel : `.claude/audits/external/`.
|
||||||
2. Collé — collez ici le contenu du PDF ou le "prompt pour IA"
|
2. Collé — collez ici le contenu du PDF ou le "prompt pour IA"
|
||||||
que l'outil suggère.
|
que l'outil suggère.
|
||||||
3. Ignorer — continuer sans. Le rapport final recommandera
|
3. Ignorer — continuer sans. Le rapport final recommandera
|
||||||
@@ -348,6 +348,15 @@ audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas,
|
|||||||
entity SEO, content shape for AI, AI visibility) — the geo-analyzer
|
entity SEO, content shape for AI, AI visibility) — the geo-analyzer
|
||||||
agent runs in parallel and owns those.
|
agent runs in parallel and owns those.
|
||||||
|
|
||||||
|
Do NOT score security headers either (CSP, HSTS, X-Frame-Options,
|
||||||
|
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
|
||||||
|
cookie flags) — `/harden` owns them and grades them 0-100 against three
|
||||||
|
external validators (`depth-matrix.md:29`). Read them, keep
|
||||||
|
`X-Robots-Tag` under indexability (it is an indexing directive, not a
|
||||||
|
security header), and declare the rest in §14 with a "run /harden" pointer
|
||||||
|
plus what you observed live. Dropping them from the score must not make
|
||||||
|
them silent.
|
||||||
|
|
||||||
FILE OWNERSHIP (authoritative, prevents parallel-edit conflicts):
|
FILE OWNERSHIP (authoritative, prevents parallel-edit conflicts):
|
||||||
- YOU OWN (read+write): sitemap.xml, image/video sitemaps, .htaccess,
|
- YOU OWN (read+write): sitemap.xml, image/video sitemaps, .htaccess,
|
||||||
meta tags (title, description, OG, Twitter, canonical, robots meta),
|
meta tags (title, description, OG, Twitter, canonical, robots meta),
|
||||||
|
|||||||
Reference in New Issue
Block a user