Merge bugfix/seo-geo-integrity-phase1 into develop
This commit is contained in:
@@ -1,5 +1,112 @@
|
||||
# TODO
|
||||
|
||||
## 2026-07-16 — PLAN seo/geo parity vs claude-seo (not started, awaiting arbitrage)
|
||||
Source: audit of github.com/AgriciDaniel/claude-seo (11.5k★, MIT, v2.2.0,
|
||||
5 mo old, 185/197 commits single author). Verdict: cherry-pick, never install
|
||||
(install.sh:49 overwrites our skills/seo/; uninstall.sh:45 glob `seo-*.md`
|
||||
deletes our seo-analyzer.md 42K it never installed; extensions/*/install.sh:42
|
||||
wipes settings.json on parse error; skills/seo/SKILL.md:119 injects Skool
|
||||
upsell footer into deliverables). Their code is real (render_page.py 428 l
|
||||
Playwright, url_safety.py 622 l SSRF, 326 tests, 320 pass) — adapt to our
|
||||
fetch.sh contract, do NOT copy wholesale (no fail-open, no tokenstore, no
|
||||
JSON shape).
|
||||
|
||||
Framing: their plus-values map onto OUR integrity gaps — report claims more
|
||||
than it measured. Same bar we held their README to.
|
||||
Seam: `lib/seo-data/fetch.sh` verbs (accounts|crux|queries|inspect|forget)
|
||||
+ fail-open `{"status":"degraded"}` + fixtures + tests. Everything below lands
|
||||
as NEW VERBS. No new architecture.
|
||||
|
||||
### AXE 0 — Integrity (no new deps, hours) — the score currently lies
|
||||
- [ ] I1 Off-page axis scores 10-15% of FULL with ZERO data source (no API,
|
||||
no index) → today fabricated, and it feeds /client-handover. Immediate
|
||||
fix: extend existing LOCAL `N/A — requires FULL audit` pattern to FULL,
|
||||
redistribute weights. Data upgrade later (AXE 3). Honesty now, data after.
|
||||
- [ ] I2 VSI (Visual Stability Index) listed in CWV thresholds but NO path
|
||||
retrieves it — neither CrUX nor PSI expose it. Phantom signal → remove
|
||||
or source.
|
||||
- [ ] I3 **SAFETY** /geo standalone: geo/SKILL.md (125 l) has no STEP 0, no
|
||||
confirmed-NAP collection — but geo-analyzer OWNS JSON-LD NAP. Standalone
|
||||
/geo on a local business can write unverified NAP with zero LRN-032
|
||||
protection. Real bug, not cosmetic.
|
||||
- [ ] I4 Security headers counted 3× (seo-analyzer STEP 4 scores them in
|
||||
Technical axis; depth-matrix.md says drop unless indexability; /harden
|
||||
re-audits /100 with 3 validators). Contradiction between dedup rule and
|
||||
agent spec → pick one owner.
|
||||
- [ ] I5 Report says "audit", measured 5-15 sampled pages. State coverage %
|
||||
explicitly in §0 until AXE 2 lands.
|
||||
|
||||
### AXE 1 — Free wins on auth we ALREADY have (fetch.sh verbs)
|
||||
- [ ] W1 `richresults` verb — GSC URL Inspection already returns
|
||||
`richResultsResult`; our OAuth already carries the scope. Programmatic
|
||||
rich-results validation on real Google data. **BEATS claude-seo**: their
|
||||
README:314 "dual validator (Rich Results Test + Markup Validator)" is
|
||||
FALSE — grep of all .py = zero calls, they are hyperlinks a human clicks.
|
||||
Today our JSON-LD validity is LLM-read only.
|
||||
- [ ] W2 `bing` verb — Bing Webmaster API, free. Closes the Google/Bing
|
||||
asymmetry (Google = full OAuth layer, Bing = manual checklist) while
|
||||
/geo targets ChatGPT Search, which indexes via Bing. Strategic, not cosmetic.
|
||||
- [ ] W3 `sameas` resolution check — trivial curl loop. entity-seo.md lists
|
||||
"sameAs pointing to dead profiles" as a known error class and never
|
||||
checks it. ~10 lines.
|
||||
|
||||
### AXE 2 — Coverage (biggest lever: ~97% of a 500-page site unseen today)
|
||||
- [ ] C1 `crawl` verb — sitemap-driven URL discovery (we ALREADY fetch
|
||||
sitemap.xml) + deterministic sampling + coverage % reported. No Chromium,
|
||||
no paid API. Turns "5-15 LLM-chosen pages" into measured coverage.
|
||||
Tradeoff vs claude-seo's link-following 500-page crawl: cheaper, but
|
||||
misses unlinked/unsitemapped pages — accept + disclose.
|
||||
- [ ] C2 Dupe/cannibalization detection — becomes possible once N pages in
|
||||
hand: compare titles/H1/canonicals across the set. Free, unblocked by C1.
|
||||
- [ ] C3 Internal-link graph — orphan pages + 3-click depth are TODAY stated
|
||||
as checks with no command to compute them. C1 unblocks real computation.
|
||||
|
||||
### AXE 3 — Off-page real (upgrades I1)
|
||||
- [ ] B1 `backlinks` verb — Common Crawl hyperlinkgraph
|
||||
(data.commoncrawl.org/projects/hyperlinkgraph), free, no key.
|
||||
- [ ] B2 Honest cap — steal their idea (free-backlink-sources.md:33: cap
|
||||
health at 70/100 when only Common Crawl). Fits our code-ceiling doctrine
|
||||
exactly.
|
||||
- [ ] B3 VERIFY FIRST: GSC Links API. Subagent claimed "available, OAuth
|
||||
already there" — I doubt it: Search Console API v3 has no links endpoint
|
||||
(links report is UI-only AFAIK). Verify before planning on it. Do not
|
||||
assert.
|
||||
|
||||
### AXE 4 — SPA blindness (dep decision — needs arbitrage)
|
||||
- [ ] R1 `render` verb — Playwright, GATED on SPA detection (STEP 2 already
|
||||
detects framework + rendering mode). Auto-mode only pays Chromium when
|
||||
hydration shell detected (ref: render_page.py:226 logic, adapt not copy).
|
||||
- [ ] R2 ARBITRAGE: heavy dep (Chromium ~300MB) vs our bash+curl purity.
|
||||
Cheaper honest alternative: on SPA, REFUSE to score on-page rather than
|
||||
score it wrong (today: curl reads source, not hydrated DOM → every
|
||||
meta/JSON-LD/heading/img grep is blind, compensated only by a §0 flag).
|
||||
|
||||
### AXE 5 — Hardening + regression (lower priority)
|
||||
- [ ] H1 SSRF guard on curl paths — both agents curl user-supplied domains.
|
||||
Our own CLAUDE.md doctrine says "never trust user input". url_safety.py
|
||||
(622 l, obfuscated-IPv4 decode, DNS pinning) is a solid reference.
|
||||
- [ ] H2 `drift` baseline (SQLite) — SEO.md Historique keeps only date+score+
|
||||
key changes. Their seo-drift is on-page regression detection, NOT rank
|
||||
tracking (common misread). Optional.
|
||||
|
||||
### NOT DOING (explicit, with reason)
|
||||
- Keyword volumes → Google Ads Tier 3 needs ACTIVE ad spend (~$150-300/mo);
|
||||
without spend the API returns buckets ("1K-10K"). Their own detect_tier()
|
||||
never even returns 3 (google_auth.py:642-724 caps at 2) + google-ads absent
|
||||
from requirements.txt. Not worth it.
|
||||
- Real AI SoV (ChatGPT/Perplexity citation tracking) → paid everywhere
|
||||
(SE Ranking/Profound/DataForSEO). Our current honest "not testable, here's
|
||||
what we measured instead" disclosure BEATS faking it. Keep.
|
||||
- Installing the plugin / +33 skills namespace → see destructive paths above.
|
||||
|
||||
### Keep (already beats claude-seo — do not regress)
|
||||
FR legal (LCEN/RGPD-ePrivacy/DGCCRF L121-1 — their whole repo: 2 hits, and
|
||||
dma-consent-mode-v2.md:27 tells the agent to stay out) · fix-bundle +
|
||||
ownership matrix + serial apply (their 18 agents are report-only, no
|
||||
ownership discipline) · trajectory-to-17/20 + honest code ceiling (theirs is
|
||||
flat 0-100, no legal axis) · llms.txt honest framing · NAP anti-dup-seed
|
||||
(LRN-032).
|
||||
|
||||
## 2026-07-16 — /close auto-persist memory (feature/close-auto-persist, BDR-068)
|
||||
- [x] STEP 5C: auto-finish chore→develop + push when capitalize/close branched off develop
|
||||
- [x] --no-push escape hatch; WORKING-branch + rc-3 skip; graceful push-fail
|
||||
|
||||
+57
-7
@@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the
|
||||
|
||||
## Context — why GEO is its own discipline in 2026
|
||||
|
||||
- AI Overviews trigger on ~48% of Google searches (April 2026).
|
||||
- ChatGPT processes 2.5B queries/day.
|
||||
- Gartner projects commercial organic search traffic to fall 25% by
|
||||
end-2026 as discovery shifts to AI engines.
|
||||
- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google
|
||||
searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner
|
||||
projects commercial organic search traffic to fall 25% by end-2026 as
|
||||
discovery shifts to AI engines. Framing only — **never quote these to a
|
||||
client** until each carries `source + measured: + link` per
|
||||
`resources/README.md`. GEO is worth doing on mechanism; it does not need
|
||||
these numbers to be true.
|
||||
- Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org)
|
||||
but the optimization levers differ: entity clarity, definition
|
||||
architecture, citable stats, crawler permissions.
|
||||
@@ -141,6 +144,16 @@ If called standalone via `/geo`, gather:
|
||||
|
||||
## STEP 2 — DETECT CONTEXT `[both]`
|
||||
|
||||
**FIRST — the CWD must BE the audited site.** You grep the current working
|
||||
directory; no dispatcher checks that it matches the target domain. If a URL
|
||||
was supplied and the CWD shows no web project at all (no `package.json` /
|
||||
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
|
||||
signals contradict the domain, STOP and report:
|
||||
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
|
||||
confirm live-only audit (LOCAL findings will be N/A).`
|
||||
Never grep one codebase while curling another: the live half looks right,
|
||||
the code half is fiction, and the report reads as authoritative.
|
||||
|
||||
```bash
|
||||
# Framework (reuse detection from seo-analyzer if available)
|
||||
ls package.json composer.json Gemfile Cargo.toml go.mod 2>/dev/null
|
||||
@@ -360,7 +373,9 @@ action (G5 batch, confirmation needed — visible page creation).
|
||||
|
||||
**Local business:**
|
||||
- [ ] `LocalBusiness` with most specific subclass (Plumber/Dentist/etc.)
|
||||
- [ ] NAP consistent with GMB
|
||||
- [ ] NAP consistent with GMB — **direction rule applies** (Data integrity:
|
||||
never pick a value from source majority; no canonical → no directional
|
||||
fix)
|
||||
- [ ] `sameAs` includes GMB URL + main social + Wikidata if applicable
|
||||
- [ ] `areaServed` lists served cities/regions
|
||||
- [ ] `openingHoursSpecification` matches reality
|
||||
@@ -447,6 +462,12 @@ Load: `~/.claude/agents/resources/content-shape-for-ai.md`
|
||||
|
||||
Sample 5-10 key pages (homepage + top service/blog pages). For each:
|
||||
|
||||
**Record the denominator.** This samples; the report says "audit". Count the
|
||||
URLs in `sitemap.xml` for the coverage ratio, and carry it into the GEO
|
||||
SCORING block. No sitemap → total UNKNOWN, say so. Content shape is the
|
||||
axis most damaged by silent sampling: it is judged per page, so a 6-page
|
||||
sample of a 300-page site says nothing about the other 294.
|
||||
|
||||
### Checks
|
||||
|
||||
1. **Definition Lead** — does the first sentence (or H1) follow
|
||||
@@ -574,6 +595,7 @@ Score each axis. Use concrete findings from STEP 2-9.
|
||||
|
||||
```
|
||||
GEO SCORING (<depth>)
|
||||
COVERAGE : <N> of <M> sitemap URLs (<P>%) | <N> pages, total UNKNOWN
|
||||
AI Crawlers Policy : XX/20 <justification>
|
||||
llms.txt : XX/20 <justification>
|
||||
Schema.org for AI : XX/20 <justification>
|
||||
@@ -584,6 +606,12 @@ AI Visibility (live) : XX/20 | N/A (LOCAL)
|
||||
GEO GLOBAL (weighted) : XX.X/20 (<depth>)
|
||||
```
|
||||
|
||||
**COVERAGE is mandatory, never omitted, never rounded up.** It bounds the
|
||||
per-page axes — Content Shape above all, and the page-level share of
|
||||
Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected:
|
||||
robots.txt and llms.txt are single files, fully read. Say which is which
|
||||
rather than letting one ratio discredit the whole report.
|
||||
|
||||
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
||||
local, 25% for national/SaaS/content.**
|
||||
|
||||
@@ -895,9 +923,31 @@ PROCHAINE ETAPE : <highest-priority>
|
||||
- **No invented entity data.** Never write a fake Wikidata QID, fake
|
||||
`sameAs` URLs, fake `knowsAbout`, fake press mentions. Unknown →
|
||||
placeholder `[À COMPLÉTER]` or omit.
|
||||
- **NAP direction rule (LRN-032).** You own JSON-LD NAP, so this binds you
|
||||
whoever called you — `/seo` passes a canonical, standalone `/geo` does
|
||||
not. NEVER infer a correct NAP value from source majority: on-site
|
||||
sources (JSON-LD, footer, settings DB, legal pages) usually descend from
|
||||
ONE seed and can all carry the same wrong value — the single diverging
|
||||
source may be the only one a human actually corrected. Direction of fix:
|
||||
- Diverging from a CONFIRMED canonical field (passed by `/seo` STEP 0)
|
||||
→ fix the diverging source.
|
||||
- Canonical UNCONFIRMED or absent (the standalone `/geo` case) → report
|
||||
the divergence WITHOUT a directional fix; escalate as a user question
|
||||
("which value is correct?") in §11.
|
||||
No G2/G6 item may write or rewrite a NAP value that no confirmed
|
||||
canonical backs — **creating** a `LocalBusiness` from scratch included:
|
||||
unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling
|
||||
on-site source.
|
||||
- **Remove deprecated schemas rather than keep broken ones.**
|
||||
- **Cite sources.** When emitting stats in the report, link
|
||||
`content-shape-for-ai.md` research citations.
|
||||
- **Cite sources, and only citable ones.** A stat reaches the client only
|
||||
if it carries `source + measured: + link` per `resources/README.md`.
|
||||
Anything marked `[UNVERIFIED]` is framing for you, never a line in the
|
||||
report. Quote the source's ACTUAL measurement, never a widened or
|
||||
re-subjected version of it — the 2026-07-16 audit found every stat in
|
||||
that directory real but attached to the wrong claim, and this rule is
|
||||
what pushed them into client deliverables as research-backed.
|
||||
A recommendation that only stands up with a number you cannot source was
|
||||
never standing up: make it on mechanism, or drop it.
|
||||
|
||||
### Process
|
||||
- **Every user action lists automation options.** Mandatory from
|
||||
|
||||
@@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current.
|
||||
|
||||
These files capture state as of 2026-04. Crawler lists, Schema.org
|
||||
deprecations, and tool landscape shift fast. Agents MUST cross-check
|
||||
via WebSearch on each run when FULL depth is selected.
|
||||
crawler lists and tool names via WebSearch on each run when FULL depth is
|
||||
selected.
|
||||
|
||||
## Citation standard (mandatory for every statistic)
|
||||
|
||||
**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO
|
||||
blogs cross-cite each other into a consensus that looks like corroboration.
|
||||
Two 2026-07-16 audits of this directory show how it fails:
|
||||
|
||||
- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in
|
||||
`seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already
|
||||
collected it. It is absent from the CrUX API metric list and from
|
||||
web.dev. WebSearch returned the echo, not the truth.
|
||||
- Every stat in this directory was real **and attached to the wrong
|
||||
subject**: the GEO paper's 40% (all methods) pinned on one technique;
|
||||
LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay;
|
||||
AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with
|
||||
its meaning inverted; a smart-speaker adoption figure sold as voice-search
|
||||
share.
|
||||
|
||||
The failure mode is not invention — it is **plausible recombination**, which
|
||||
is exactly what a model half-remembering a search result produces. So the
|
||||
format has to make an unsourced number conspicuous:
|
||||
|
||||
```
|
||||
<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
|
||||
measured> — <link>
|
||||
```
|
||||
|
||||
`measured:` is the field that catches it. All four errors above survive a
|
||||
source name; none survives having to state the source's real measurement
|
||||
next to the claim.
|
||||
|
||||
Rules:
|
||||
1. **Primary source or no number.** Peer-reviewed paper, the vendor's own
|
||||
published study, or an official API/doc. `developer.chrome.com/docs/crux`
|
||||
is decisive for metrics: what CrUX cannot return, we cannot score.
|
||||
2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast,
|
||||
Ahrefs publish useful data and sell products — say "vendor".
|
||||
3. **Never widen scope.** An aggregate result is not a per-technique result.
|
||||
4. **No number beats a wrong number.** A recommendation that only stands up
|
||||
with a fabricated statistic was never standing up. Delete the stat, keep
|
||||
the recommendation if it survives on mechanism.
|
||||
5. **Unverified ⇒ labelled.** `[UNVERIFIED — <date>]` inline. Never quote an
|
||||
unverified number to a client: `geo-analyzer.md` ("Cite sources") sends
|
||||
these into client reports as research-backed.
|
||||
|
||||
## Loading pattern
|
||||
|
||||
|
||||
@@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers
|
||||
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
|
||||
Overviews.
|
||||
|
||||
Context: Google AI Overviews trigger on ~48% of searches; ChatGPT
|
||||
processes 2.5B queries/day; Gartner projects commercial organic
|
||||
search traffic will drop 25% by 2026. Monitoring is no longer optional.
|
||||
Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of
|
||||
searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial
|
||||
organic search traffic will drop 25% by 2026.
|
||||
|
||||
> Not checked against primary sources in the 2026-07-16 audit that corrected
|
||||
> the rest of this directory — flagged rather than asserted or deleted, per
|
||||
> the citation standard in `README.md` (rule 5). The Gartner projection at
|
||||
> least names its source; the other two float. Treat all three as
|
||||
> motivation, not evidence: **do NOT quote them to a client** until each
|
||||
> carries `source + measured: + link`. Their only job here is to explain why
|
||||
> this file exists, and that argument does not need numbers.
|
||||
|
||||
## Commercial tools
|
||||
|
||||
|
||||
@@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density.
|
||||
|
||||
### 4. Citations and statistics (strongest measured lever)
|
||||
|
||||
Adding peer-cited statistics with clear sources increases AI visibility
|
||||
**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine
|
||||
Optimization").
|
||||
Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024)
|
||||
report that their optimisation methods **collectively** boost visibility
|
||||
**by up to 40%** in generative-engine responses, and state the effect
|
||||
**varies across domains**. Citations/statistics/quotations are among those
|
||||
methods.
|
||||
|
||||
> **Attribute this correctly.** Until 2026-07-16 this section read "Adding
|
||||
> peer-cited statistics with clear sources increases AI visibility by up to
|
||||
> 40%" — pinning the paper's *aggregate* result on this *one* technique. The
|
||||
> paper publishes no separate figure per technique. When quoting it to a
|
||||
> client: "up to 40%, across the method set, domain-dependent" — never "+40%
|
||||
> if you add stats".
|
||||
|
||||
Pattern: embed specific numbers with attribution.
|
||||
|
||||
@@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure:
|
||||
|
||||
### 6. Freshness signals
|
||||
|
||||
Pages not updated at least quarterly are **3x more likely to lose AI
|
||||
citations** (LLMRefs 2026 study).
|
||||
Freshness is a real retrieval input: RAG systems fetch live and read
|
||||
timestamps, so a page updated this quarter carries a stronger recency
|
||||
signal than the same page last touched years ago. LLMrefs (a **vendor**,
|
||||
not peer review) reports cited content running **~25.7% fresher** than
|
||||
organic top-10 across ~17M citations. Substantive updates only — bumping a
|
||||
date string is not freshness.
|
||||
|
||||
> **The "3x" that lived here was grafted from another claim.** Until
|
||||
> 2026-07-16 this read "Pages not updated at least quarterly are 3x more
|
||||
> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x"
|
||||
> says **brand mentions correlate ~3x more strongly with AI visibility than
|
||||
> backlinks** — a different subject entirely. No source supports a quarterly
|
||||
> decay multiplier. Recommend quarterly refresh on its merits; do not price
|
||||
> it with a borrowed number.
|
||||
|
||||
What to maintain:
|
||||
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
|
||||
|
||||
@@ -21,8 +21,20 @@ existing instances. They no longer produce rich results.
|
||||
|
||||
### QAPage — single Q&A format
|
||||
|
||||
Pages cited 58% more often by ChatGPT vs basic Article schema.
|
||||
Use when the page is built around ONE primary question.
|
||||
Use when the page is built around ONE primary question. Emitting the type
|
||||
that matches the content shape beats wrapping everything in a generic
|
||||
`Article`.
|
||||
|
||||
> **No lift figure here — the one that lived here was wrong.** Until
|
||||
> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic
|
||||
> Article schema", uncited. Nothing supports it. The nearest real number is
|
||||
> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews /
|
||||
> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%**
|
||||
> of cited sources — a *prevalence* count for a *different type* — while
|
||||
> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim
|
||||
> it was propping up. Q&A shape is still worth doing on genuinely
|
||||
> single-question pages; it is not worth a fabricated number. Do NOT quote a
|
||||
> QAPage lift % to a client — there isn't one.
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -81,8 +93,16 @@ visible content.
|
||||
|
||||
### Speakable — voice + AI extraction marker
|
||||
|
||||
62% of searches in 2026 involve voice. Speakable flags the passage
|
||||
best suited for voice readout and AI summary.
|
||||
Speakable flags the passage best suited for voice readout and AI summary.
|
||||
|
||||
> **No voice-share figure — the one that lived here was a conflation.**
|
||||
> Until 2026-07-16 this read "62% of searches in 2026 involve voice",
|
||||
> uncited. No primary source carries it; 62% circulates as a *smart-speaker
|
||||
> adoption* number, not a share of searches. It is the same family as the
|
||||
> "50% of searches will be voice by 2020" myth — attributed to ComScore,
|
||||
> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable
|
||||
> is cheap and harmless, so keep recommending it on TL;DR / summary blocks —
|
||||
> but justify it by extraction shape, never by a voice-share statistic.
|
||||
|
||||
```json
|
||||
{
|
||||
|
||||
+121
-6
@@ -81,6 +81,17 @@ hreflang, infer from detected URL structures.
|
||||
|
||||
## STEP 2 — DETECT TECHNICAL CONTEXT `[both]`
|
||||
|
||||
**FIRST — the CWD must BE the audited site.** You grep the current working
|
||||
directory; no dispatcher checks that it matches TARGET_URL. If a URL was
|
||||
supplied and the CWD shows no web project at all (no `package.json` /
|
||||
`composer.json` / `index.html` / `*.astro` / `*.php` / `.htaccess`), or its
|
||||
signals contradict the domain, STOP and report:
|
||||
`CWD/TARGET MISMATCH — <cwd> is not <domain>'s repo. Re-run from it, or
|
||||
confirm live-only audit (LOCAL findings will be N/A).`
|
||||
Never grep one codebase while curling another: the live half looks right,
|
||||
the code half is fiction, and the report reads as authoritative. `/harden`
|
||||
inherits this agent for its config axis, so the mismatch propagates there.
|
||||
|
||||
### Framework & rendering
|
||||
|
||||
```bash
|
||||
@@ -148,6 +159,23 @@ RECOMMENDATION : KEEP & CONFIGURE plugin | INSTALL <plugin> (P0 quick win) | M
|
||||
|
||||
### Infrastructure signals
|
||||
|
||||
**Origin vs edge — never infer the stack from `server:`.** That header names
|
||||
whatever answered: usually the EDGE (Cloudflare, Scaleway/OVH front, CDN,
|
||||
load balancer), not the origin. Apache behind an nginx front is a standard
|
||||
topology — TLS terminated upstream, the origin sees plain HTTP plus
|
||||
`X-Forwarded-Proto`.
|
||||
- Repo `.htaccess` + `server: nginx` = NOT drift, NOT dead config. Do not
|
||||
flag it, do not propose migrating it.
|
||||
- Never move headers into an `nginx.conf` absent from the repo. Server-side
|
||||
config you cannot read is a §14 gap, not a finding.
|
||||
- A header present live but in no repo config = "set upstream", never
|
||||
"missing".
|
||||
|
||||
`/harden` reuses this agent for its entire config-hardening axis, so a wrong
|
||||
topology call scores a client's server config against a file that never ran.
|
||||
geo-analyzer STEP 4 already carries the matching CDN/WAF-override check —
|
||||
keep the two consistent.
|
||||
|
||||
```bash
|
||||
# Server / hosting
|
||||
ls .htaccess nginx.conf netlify.toml vercel.json wrangler.toml 2>/dev/null
|
||||
@@ -216,6 +244,12 @@ anonymous PageSpeed lab data and STEP 4/STEP 11 emit the §11 user action
|
||||
|
||||
### HTTP headers & security
|
||||
|
||||
**Read them; score them only for `/harden` (I4).** This section stays — the
|
||||
raw headers are needed for `X-Robots-Tag`, canonical/redirect coherence, and
|
||||
the §14 observed-list. But under `/seo` the security headers themselves are
|
||||
out of scope for scoring: see the Technical axis note in STEP 9. Under
|
||||
`/harden` they are the entire job. Reading is not scoring.
|
||||
|
||||
```bash
|
||||
DOMAIN="<production-domain>"
|
||||
|
||||
@@ -247,8 +281,21 @@ Evaluate each present/missing:
|
||||
- **LCP** (Largest Contentful Paint) — < 2.5s
|
||||
- **INP** (Interaction to Next Paint) — < 200ms (replaced FID in Mar 2024)
|
||||
- **CLS** (Cumulative Layout Shift) — < 0.1
|
||||
- **VSI** (Visual Stability Index) — new 2026 signal, Google Core Web
|
||||
Vitals 2.0
|
||||
|
||||
**Core Web Vitals are exactly these three** (web.dev/articles/vitals,
|
||||
verified 2026-07-16). Google ships threshold changes with prior notice on a
|
||||
predictable annual cadence — a "new CWV" that only SEO blogs know about does
|
||||
not exist. Before adding a metric here, confirm it against a PRIMARY source:
|
||||
web.dev, the Chromium blog, or `developer.chrome.com/docs/crux/api` — that
|
||||
API metric list is decisive, because a metric CrUX cannot return is a metric
|
||||
we cannot score.
|
||||
|
||||
**WebSearch is not confirmation.** SEO blogs cross-cite each other into fake
|
||||
consensus. A "VSI (Visual Stability Index) — new 2026 signal, Core Web
|
||||
Vitals 2.0" line lived here until 2026-07-16 on exactly that basis: ten
|
||||
blogs asserted it, several claimed CrUX was already collecting it, and it is
|
||||
absent from both the CrUX API metric list and web.dev. Stated as fact, in a
|
||||
threshold list, in client-facing audits.
|
||||
|
||||
When a GSC account+property were passed in context, fetch CrUX field
|
||||
data first (**tilde path mandatory** — this agent runs from the
|
||||
@@ -359,8 +406,22 @@ Fetch rendered HTML. Extract and analyze:
|
||||
|
||||
## STEP 5 — ON-PAGE AUDIT `[both]`
|
||||
|
||||
**Record the denominator BEFORE sampling.** This step samples; the report
|
||||
says "audit". Count the URLs in `sitemap.xml` (fetch it in full — the
|
||||
`head -50` in STEP 4 is a preview, not a count). That count is the coverage
|
||||
denominator, and it feeds the mandatory COVERAGE line in STEP 9. No sitemap
|
||||
→ denominator unknown: say so, never let silence imply full coverage. On a
|
||||
500-page site a 12-page sample is 2.4% — the On-page score is an
|
||||
extrapolation from it, and the reader cannot know that unless you print it.
|
||||
|
||||
### Meta tags per page (sample 5-15 key pages)
|
||||
|
||||
Sample by risk, not convenience: homepage + top templates (one per page
|
||||
type: service, city, blog, product, legal) + any page GSC flags as a
|
||||
position 4-10 quick win. Same template audited twice buys nothing; an
|
||||
un-sampled template is an un-audited template — name the templates you
|
||||
skipped.
|
||||
|
||||
For each sampled page:
|
||||
```
|
||||
PAGE: <path>
|
||||
@@ -616,10 +677,10 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
||||
|
||||
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
||||
|---|---|---|---|
|
||||
| Technical (perf, CWV, security headers, indexability) | 20% | 30% | |
|
||||
| Technical (perf, CWV, indexability) | 20% | 30% | |
|
||||
| On-page (content, meta, headings, images, video, a11y, i18n) | 20% | 30% | |
|
||||
| SEO Local (NAP, GMB, citations) | 25% | 5% | |
|
||||
| Off-page (backlinks, mentions, authority) | 10% | 15% | |
|
||||
| Off-page (unlinked brand mentions — backlinks/authority NOT auditable, §14) | 10% | 15% | |
|
||||
| Social presence | 10% | 5% | |
|
||||
| Competitive position | 5% | 10% | |
|
||||
| Legal compliance | 10% | 5% | |
|
||||
@@ -628,17 +689,63 @@ FIX: AUTO (<what agent will do>) | USER (<what user must do>)
|
||||
real users, from STEP 4) when available; otherwise lab PageSpeed
|
||||
Lighthouse run.
|
||||
|
||||
**Security headers are NOT scored here (I4).** `/harden` owns them and
|
||||
grades them out of 100 with three external validators — pricing them into
|
||||
this axis too was double-counting the same finding in two reports
|
||||
(`depth-matrix.md:29` already said drop; this spec contradicted it).
|
||||
- Dispatched from `/harden` (its prompt says NARROW-SCOPE): headers ARE the
|
||||
job — audit and score them per its brief, ignore this note.
|
||||
- Dispatched from `/seo`: do not score CSP, HSTS, X-Frame-Options,
|
||||
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
|
||||
cookie flags. STEP 4 still reads them — you need them for the one
|
||||
carve-out below — but they earn and lose no points here.
|
||||
|
||||
**Carve-out — `X-Robots-Tag` stays.** It is an indexing directive wearing a
|
||||
header's clothes: `noindex` served there deindexes the page as surely as a
|
||||
meta robots tag. Score it under indexability. That is what
|
||||
`depth-matrix.md:29` means by "unless it directly affects indexability" —
|
||||
it is the header that does, and the security headers above are not.
|
||||
|
||||
**Drop ≠ silence.** A user who never runs `/harden` must not read a clean
|
||||
Technical score as clean headers. Whenever depth=FULL, emit in §14:
|
||||
`Security headers (CSP, HSTS, X-Frame-Options…) — not scored here: /harden
|
||||
owns them (0-100 + Observatory/SecurityHeaders/SSL Labs). Run /harden
|
||||
<url>. Observed live this run: <present list | none observed>.`
|
||||
Name what you saw. An omission has to stay legible — the same reason
|
||||
COVERAGE is mandatory in STEP 9.
|
||||
|
||||
**Off-page axis note (I1).** Score ONLY the unlinked brand mentions
|
||||
gathered in STEP 6 (`web_search "<business-name>" -site:<domain>`).
|
||||
Backlink profile and domain authority have NO data source here — no index,
|
||||
no API, nothing. NEVER price them into the number: an unmeasured
|
||||
sub-component cannot be judged, and this axis carries 10-15% of a score
|
||||
that reaches a client via `/client-handover`. A low mention count is a low
|
||||
mention count — it is NOT evidence of a weak backlink profile.
|
||||
|
||||
Mandatory §14 line whenever depth=FULL, verbatim:
|
||||
`Backlinks / domain authority — NOT audited: no backlink index wired.
|
||||
Nearest free source: Common Crawl hyperlinkgraph. Commercial: Ahrefs /
|
||||
Semrush / Majestic. The Off-page score above prices in brand mentions only.`
|
||||
|
||||
Weight deliberately unchanged despite the narrower scope: re-deriving it
|
||||
now, then again when a backlink source lands, would churn historical
|
||||
scores twice. Revisit the 10/15% only when the axis widens back.
|
||||
|
||||
### LOCAL depth — 4 axes
|
||||
|
||||
| Axis | Weight (local B2C) | Weight (SaaS/national/content) | Score /20 |
|
||||
|---|---|---|---|
|
||||
| Technical (security headers, indexability, config) | 25% | 35% | |
|
||||
| Technical (indexability, config) | 25% | 35% | |
|
||||
| On-page (content, meta, headings, images, video, a11y, i18n) | 35% | 45% | |
|
||||
| SEO Local (markup, NAP in JSON-LD, legal) | 20% | 5% | |
|
||||
| Legal compliance (pages, CMP, mentions) | 20% | 15% | |
|
||||
|
||||
LOCAL axes not audited (Off-page, Social, Competitive) appear as
|
||||
`N/A — requires FULL audit` in the report.
|
||||
`N/A — requires FULL audit` in the report. Off-page is the exception to
|
||||
that promise: FULL audits its brand-mentions share ONLY — backlinks and
|
||||
authority are unauditable at EVERY depth (see the Off-page axis note).
|
||||
Print `N/A — FULL audits brand mentions only` for it, never a bare
|
||||
"requires FULL audit" that FULL cannot keep.
|
||||
|
||||
### Projected code-only score + trajectory to 17/20 (mandatory)
|
||||
|
||||
@@ -676,6 +783,8 @@ misroutes the client-handover gate and the user's effort.
|
||||
|
||||
```
|
||||
SEO SCORING (<depth>)
|
||||
COVERAGE : <N> of <M> sitemap URLs (<P>%) — templates skipped: <list|none>
|
||||
| <N> pages, total UNKNOWN (no sitemap)
|
||||
Technical : XX/20 <justification>
|
||||
On-page : XX/20 <justification>
|
||||
SEO Local : XX/20 | N/A
|
||||
@@ -687,6 +796,12 @@ Legal : XX/20 <justification>
|
||||
SEO GLOBAL (weighted): XX.X/20 (<depth>)
|
||||
```
|
||||
|
||||
**COVERAGE is mandatory, never omitted, never rounded up.** It is the
|
||||
honesty bound on every page-level axis: On-page and the on-page share of
|
||||
Technical are extrapolations from the sample. If coverage < 25%, repeat it
|
||||
in §0 as a major alert — a 17/20 drawn from 3% of a site is not a 17/20, and
|
||||
`/client-handover` gates on these numbers.
|
||||
|
||||
Per user instruction: this score represents **80% of the combined
|
||||
final score for local B2C (20% for GEO), or 75% for SaaS/national
|
||||
(25% for GEO)**. The `/seo` dispatcher combines SEO and GEO scores.
|
||||
|
||||
+12
-3
@@ -204,9 +204,9 @@ Ask ONCE before dispatching the agents:
|
||||
```
|
||||
RAPPORT EXTERNE (optionnel) — un autre regard sur le site :
|
||||
|
||||
1. Fichier — déposez l'export (PDF/MD/TXT) dans
|
||||
`.claude/audits/external/` (ex. `sorank-YYYY-MM-DD.pdf`),
|
||||
donnez le nom du fichier. (`mkdir -p .claude/audits/external`)
|
||||
1. Fichier — donnez le chemin de l'export (PDF/MD/TXT), où qu'il soit
|
||||
(ex. `~/Téléchargements/sorank-2026-07-16.pdf`). Rangement conseillé
|
||||
mais optionnel : `.claude/audits/external/`.
|
||||
2. Collé — collez ici le contenu du PDF ou le "prompt pour IA"
|
||||
que l'outil suggère.
|
||||
3. Ignorer — continuer sans. Le rapport final recommandera
|
||||
@@ -348,6 +348,15 @@ audit GEO/AI signals (llms.txt, AI crawlers, QAPage/Speakable schemas,
|
||||
entity SEO, content shape for AI, AI visibility) — the geo-analyzer
|
||||
agent runs in parallel and owns those.
|
||||
|
||||
Do NOT score security headers either (CSP, HSTS, X-Frame-Options,
|
||||
X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/CORP,
|
||||
cookie flags) — `/harden` owns them and grades them 0-100 against three
|
||||
external validators (`depth-matrix.md:29`). Read them, keep
|
||||
`X-Robots-Tag` under indexability (it is an indexing directive, not a
|
||||
security header), and declare the rest in §14 with a "run /harden" pointer
|
||||
plus what you observed live. Dropping them from the score must not make
|
||||
them silent.
|
||||
|
||||
FILE OWNERSHIP (authoritative, prevents parallel-edit conflicts):
|
||||
- YOU OWN (read+write): sitemap.xml, image/video sitemaps, .htaccess,
|
||||
meta tags (title, description, OG, Twitter, canonical, robots meta),
|
||||
|
||||
Reference in New Issue
Block a user